Ahsen Tahir

Publications

Conference & workshop papers

Anemoi: A Semi-Centralized Multi-agent System Based on Agent-to-Agent Communication MCP server from Coral Protocol Xinxing Ren, Caelum Forder, Qianbo Zang, Ahsen Tahir, Roman J. Georgio, Suman Deb, Peter Carroll, Önder Gürcan, Zekun Guo NeurIPS 2025Workshop on Language Agents and World Models (LAW)

A semi-centralized multi-agent system that replaces a single all-knowing planner with structured agent-to-agent communication, reaching 52.73% on GAIA with a small planner model and beating the OWL baseline by +9.09% under identical LLM settings.

arXiv:2508.17068·BibTeX·details

Preprints

Beyond Rule-Based Workflows: An Information-Flow-Orchestrated Multi-Agents Paradigm via Agent-to-Agent Communication from CORAL Xinxing Ren, Qianbo Zang, Caelum Forder, Suman Deb, Ahsen Tahir, Roman J. Georgio, Peter Carroll, Zekun Guo arXiv preprint arXiv:2601.09883, January 2026

A workflow-free multi-agent paradigm in which a centralized information-flow orchestrator coordinates agents communicating in natural language, reaching 63.64% pass@1 on GAIA, +8.49% over a workflow-based baseline at comparable token cost.

arXiv:2601.09883·BibTeX·details

In progress

SASD: Source-Aware Self-Distillation Prompt-injection resistance trained into the weights, not bolted on at inference In progress · targeting NeurIPS / ICLR

An agent reads retrieved pages, tool output and its operator's instructions through one undifferentiated channel, with nothing recording which span came from where. Existing defenses patch this at inference (a detector in front of the model, or delimiters around the untrusted region), and both give way to an attacker who rewords the payload.

SASD moves this into training. The student is optimized against a multi-teacher KL divergence objective where each teacher sees the same task under a different source condition, targeting the one for which the injected span was never present. Source tracking ends up in the weights, instead of a wrapper policing the model at run time.

Reproduced and benchmarked existing prompt-injection defenses on AgentDojo, establishing baselines across prior methods; the proposed approach already beats them, with ongoing experiments extending to privacy leakage and calibrated refusal.

Work in progress. I am implementing the training stack and the injection benchmark suite. Happy to discuss the work directly.

  • Prompt injection
  • Knowledge distillation
  • AI safety
Detecting Design-Level Security Anti-Patterns in LLM-Generated Microservices Ahsen Tahir, Francis Palma, Shazil Khan Ongoing · University of New Brunswick

LLM-generated code security is often evaluated at the level of line-level weaknesses such as CWEs inside individual functions. Our project asks a different question: do coding agents introduce design-level security anti-patterns when they generate whole microservice systems?

We first build a literature-grounded catalog of nine microservice security anti-patterns, each with explicit detection criteria and not-applicable conditions. We then build an agentic detector that packs and analyzes a repository, generates an architecture diagram, reconstructs service inventory, inter-service communication graph, deployment topology and auth/authz surfaces, and returns a present / absent / not-applicable verdict with evidence and rationale for each anti-pattern.

Validation is ongoing on 20 developer-built repositories using two independent human raters, with a third reviewer adjudicating disagreements. This human-validated set is the benchmark for testing the detector itself across coding agents including Claude, OpenAI Codex, Google Antigravity CLI and Microsoft Copilot. Next, we will generate new microservice repositories with these agents and use the validated pipeline to study whether LLM-generated architectures exhibit design-level security anti-patterns.

  • Software security
  • LLM-generated code
  • Microservices
  • Agentic evaluation