Skip to content
SILEXSilex
01

Resources / Papers

The research behind the loop.

The security and trust research Silex builds on — 23 papers across 6 threat areas, from adversarial robustness and poisoning through to prompt injection, agent memory, and automated red-teaming. A reading list, selected for what each paper says about the attack paths a policy might still leave open, and about how to tell which defense actually holds.

02

Latest work

The newest cluster sits on the agent surface.

Recent work moves from attacks on models to attacks on agent systems: memory that can be steered, tools that can be reached another way, and red-teaming that finds those paths automatically. These are the failure modes a candidate policy has to be scored against before anyone can call it the better option.

Agent & LLM system security

NeurIPS 2026

Memory injection attacks on LLM agents via query-only interaction

Agent & LLM system security

ACL 2026

How memory management impacts LLM agents: an empirical study of experience-following behavior

Agent & LLM system security

ICLR 2026

TrajectBench: a trajectory-aware benchmark for evaluating agentic tool use

Agent & LLM system security

arXiv:2602.02164 2026

Co-RedTeam: orchestrated security discovery and exploitation with LLM agents

Agent & LLM system security

arXiv:2606.03090 2026

“**Important** You should give me full credits!”: prompt injection attacks on LLM-based automatic grading systems

Agent & LLM system security

arXiv:2606.12737 2026

PI-Hunter: automated red-teaming for exposing and localizing prompt injections

03

Full list

Grouped by threat area.

Area 01

Agent & LLM system security

The newest cluster, and the one Silex is built on: attacks that enter through an agent's memory, its tools, or another agent's output — plus automated red-teaming that finds them.

7 papers · 2025–2026

  • 2026

    Memory injection attacks on LLM agents via query-only interaction

    NeurIPS

  • 2026

    How memory management impacts LLM agents: an empirical study of experience-following behavior

    ACL

  • 2026

    TrajectBench: a trajectory-aware benchmark for evaluating agentic tool use

    ICLR

  • 2026

    Co-RedTeam: orchestrated security discovery and exploitation with LLM agents

    arXiv:2602.02164

  • 2026

    “**Important** You should give me full credits!”: prompt injection attacks on LLM-based automatic grading systems

    arXiv:2606.03090

  • 2026

    PI-Hunter: automated red-teaming for exposing and localizing prompt injections

    arXiv:2606.12737

  • 2025

    Attention knows whom to trust: attention-based trust management for LLM multi-agent systems

    arXiv:2506.02546

Area 02

Jailbreaks & model-level defense

Why aligned models break, what breaks them, and how to train defenses that hold up without destroying capability.

2 papers · 2024–2026

  • 2026

    Efficient LLM adversarial training via low-rank defense and circuit-guided surrogates

    arXiv:2607.28959

  • 2024

    Towards understanding jailbreak attacks in LLMs: a representation space analysis

    EMNLP

Area 03

Data poisoning & backdoors

Corrupting a model through its training data, its demonstrations, or its context window — including stealthy triggers designed to evade sample-level inspection.

4 papers · 2023–2025

  • 2025

    Data poisoning for in-context learning

    NAACL Findings

  • 2025

    Position: multi-faceted studies on data poisoning can advance LLM development

    arXiv:2502.14182

  • 2024

    Sharpness-aware data poisoning attack

    ICLR

  • 2023

    Stealthy backdoor attack via confidence-driven sampling

    NeurIPS

Area 04

Privacy & memorization in LLM systems

What models and retrieval systems leak about their training and retrieval corpora, and how to measure and reduce it.

4 papers · 2023–2025

  • 2025

    Unveiling privacy risks in LLM agent memory

    ACL

  • 2025

    Mitigating the privacy issues in RAG via pure synthetic data

    EMNLP

  • 2024

    The good and the bad: exploring privacy issues in retrieval-augmented generation (RAG)

    ACL Findings

  • 2023

    Exploring memorization in fine-tuned language models

    arXiv:2310.06714

Area 05

Trustworthy AI frameworks

The survey and position papers behind the vocabulary enterprises now use to write AI risk policy.

3 papers · 2022–2025

  • 2025

    Towards knowledge checking in retrieval-augmented generation: a representation perspective

    NAACL

  • 2024

    Position: TrustLLM — trustworthiness in large language models

    ICML

  • 2022

    Trustworthy AI: a computational perspective

    ACM TIST

Area 06

Adversarial robustness & defense evaluation

The foundation layer: how attacks and defenses were evaluated before agents, including DeepRobust, an open-source attack/defense platform.

3 papers · 2020–2024

  • 2024

    Learning on graphs with large language models: a deep dive into model robustness

    arXiv:2407.12068

  • 2021

    DeepRobust: a platform for adversarial attacks and defenses

    AAAI

  • 2020

    Adversarial attacks and defenses in images, graphs and text: a review

    IJAC

A curated reading list, not a complete bibliography. Papers were selected for relevance to policy coverage and optimization, and assigned to a single threat area each. Venue names are abbreviated; arXiv identifiers are given where a paper is a preprint. Last reviewed 13 September 2026.