2013
Trust and adversarial behavior in online systems
Early research into detecting social spammers and understanding trust in large-scale digital systems.
Company / Team
Silex brings together research in adversarial machine learning, trustworthy AI, and agent security to help enterprises find coverage gaps and optimize the policies that close them.
Research to product
A blocked attack proves that one route was covered. Research on adaptive attacks, compromised context, and agent memory helps us ask the harder question: what other routes still reach the same unsafe outcome?
That research sensibility shapes Silex: explore alternatives, evaluate them against the real environment and business constraints, then validate candidates through simulation, shadow, and canary before production.
Research timeline
2013
Early research into detecting social spammers and understanding trust in large-scale digital systems.
2020
Research spanning attacks and defenses in images, graphs, and text established a broad foundation for evaluating adaptive threats.
2021
A platform for adversarial attacks and defenses, built to make rigorous evaluation more accessible across machine-learning systems.
2024
TrustLLM brought a structured lens to the trustworthiness questions surrounding modern language models.
2025
Research explored data poisoning for in-context learning and the privacy risks introduced when LLM agents retain memory.
2026
Work on query-only memory injection and automated prompt-injection localization brings adversarial research to agent systems.
Selected research
A selection of work that informs how we think about agent attack paths, policy evaluation, and governed defenses.
Agent security
NeurIPS 2026
How attackers can steer an agent through its memory, even when interaction is limited to queries.
Agent security
arXiv:2606.12737, 2026
Automated discovery and localization of prompt-injection failure modes in LLM systems.
Defense evaluation
AAAI 2021
A platform for rigorously evaluating adversarial attacks and defenses across machine-learning systems.
Poisoning and trust
NAACL Findings 2025
Research into how an attacker can corrupt the context a model relies on to make decisions.
Trustworthy AI
ICML 2024
A framework for understanding the trustworthiness questions that surround modern language models.
Privacy
ACL 2025
An examination of the privacy risks that arise when LLM agents retain and use memory.
Build with us
We are building the intelligence layer for continuous policy coverage and optimization in enterprise AI systems.