Agent & LLM system security
NeurIPS 2026
Resources / Papers
The security and trust research Silex builds on — 23 papers across 6 threat areas, from adversarial robustness and poisoning through to prompt injection, agent memory, and automated red-teaming. A reading list, selected for what each paper says about the attack paths a policy might still leave open, and about how to tell which defense actually holds.
Latest work
Recent work moves from attacks on models to attacks on agent systems: memory that can be steered, tools that can be reached another way, and red-teaming that finds those paths automatically. These are the failure modes a candidate policy has to be scored against before anyone can call it the better option.
Agent & LLM system security
NeurIPS 2026
Agent & LLM system security
ACL 2026
Agent & LLM system security
ICLR 2026
Agent & LLM system security
arXiv:2602.02164 2026
Agent & LLM system security
arXiv:2606.03090 2026
Agent & LLM system security
arXiv:2606.12737 2026
Full list
Area 01
The newest cluster, and the one Silex is built on: attacks that enter through an agent's memory, its tools, or another agent's output — plus automated red-teaming that finds them.
7 papers · 2025–2026
2026
Memory injection attacks on LLM agents via query-only interaction
NeurIPS
2026
How memory management impacts LLM agents: an empirical study of experience-following behavior
ACL
2026
TrajectBench: a trajectory-aware benchmark for evaluating agentic tool use
ICLR
2026
Co-RedTeam: orchestrated security discovery and exploitation with LLM agents
arXiv:2602.02164
2026
“**Important** You should give me full credits!”: prompt injection attacks on LLM-based automatic grading systems
arXiv:2606.03090
2026
PI-Hunter: automated red-teaming for exposing and localizing prompt injections
arXiv:2606.12737
2025
Attention knows whom to trust: attention-based trust management for LLM multi-agent systems
arXiv:2506.02546
Area 02
Why aligned models break, what breaks them, and how to train defenses that hold up without destroying capability.
2 papers · 2024–2026
2026
Efficient LLM adversarial training via low-rank defense and circuit-guided surrogates
arXiv:2607.28959
2024
Towards understanding jailbreak attacks in LLMs: a representation space analysis
EMNLP
Area 03
Corrupting a model through its training data, its demonstrations, or its context window — including stealthy triggers designed to evade sample-level inspection.
4 papers · 2023–2025
2025
Data poisoning for in-context learning
NAACL Findings
2025
Position: multi-faceted studies on data poisoning can advance LLM development
arXiv:2502.14182
2024
Sharpness-aware data poisoning attack
ICLR
2023
Stealthy backdoor attack via confidence-driven sampling
NeurIPS
Area 04
What models and retrieval systems leak about their training and retrieval corpora, and how to measure and reduce it.
4 papers · 2023–2025
2025
Unveiling privacy risks in LLM agent memory
ACL
2025
Mitigating the privacy issues in RAG via pure synthetic data
EMNLP
2024
The good and the bad: exploring privacy issues in retrieval-augmented generation (RAG)
ACL Findings
2023
Exploring memorization in fine-tuned language models
arXiv:2310.06714
Area 05
The survey and position papers behind the vocabulary enterprises now use to write AI risk policy.
3 papers · 2022–2025
2025
Towards knowledge checking in retrieval-augmented generation: a representation perspective
NAACL
2024
Position: TrustLLM — trustworthiness in large language models
ICML
2022
Trustworthy AI: a computational perspective
ACM TIST
Area 06
The foundation layer: how attacks and defenses were evaluated before agents, including DeepRobust, an open-source attack/defense platform.
3 papers · 2020–2024
2024
Learning on graphs with large language models: a deep dive into model robustness
arXiv:2407.12068
2021
DeepRobust: a platform for adversarial attacks and defenses
AAAI
2020
Adversarial attacks and defenses in images, graphs and text: a review
IJAC
A curated reading list, not a complete bibliography. Papers were selected for relevance to policy coverage and optimization, and assigned to a single threat area each. Venue names are abbreviated; arXiv identifiers are given where a paper is a preprint. Last reviewed 13 September 2026.