Publications
Conference & workshop papers
A semi-centralized multi-agent system that replaces a single all-knowing planner with structured agent-to-agent communication, reaching 52.73% on GAIA with a small planner model and beating the OWL baseline by +9.09% under identical LLM settings.
arXiv:2508.17068·BibTeX·detailsPreprints
A workflow-free multi-agent paradigm in which a centralized information-flow orchestrator coordinates agents communicating in natural language, reaching 63.64% pass@1 on GAIA, +8.49% over a workflow-based baseline at comparable token cost.
arXiv:2601.09883·BibTeX·detailsIn progress
An agent reads retrieved pages, tool output and its operator's instructions through one undifferentiated channel, with nothing recording which span came from where. Existing defenses patch this at inference (a detector in front of the model, or delimiters around the untrusted region), and both give way to an attacker who rewords the payload.
SASD moves this into training. The student is optimized against a multi-teacher KL divergence objective where each teacher sees the same task under a different source condition, targeting the one for which the injected span was never present. Source tracking ends up in the weights, instead of a wrapper policing the model at run time.
Reproduced and benchmarked existing prompt-injection defenses on AgentDojo, establishing baselines across prior methods; the proposed approach already beats them, with ongoing experiments extending to privacy leakage and calibrated refusal.
Work in progress. I am implementing the training stack and the injection benchmark suite. Happy to discuss the work directly.
LLM-generated code security is often evaluated at the level of line-level weaknesses such as CWEs inside individual functions. Our project asks a different question: do coding agents introduce design-level security anti-patterns when they generate whole microservice systems?
We first build a literature-grounded catalog of nine microservice security anti-patterns, each with explicit detection criteria and not-applicable conditions. We then build an agentic detector that packs and analyzes a repository, generates an architecture diagram, reconstructs service inventory, inter-service communication graph, deployment topology and auth/authz surfaces, and returns a present / absent / not-applicable verdict with evidence and rationale for each anti-pattern.
Validation is ongoing on 20 developer-built repositories using two independent human raters, with a third reviewer adjudicating disagreements. This human-validated set is the benchmark for testing the detector itself across coding agents including Claude, OpenAI Codex, Google Antigravity CLI and Microsoft Copilot. Next, we will generate new microservice repositories with these agents and use the validated pipeline to study whether LLM-generated architectures exhibit design-level security anti-patterns.