Skip to content
turenlabsPublic

About

A living catalog of LLM security research, updated daily from arXiv

Topics

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

SecPapers

A living, searchable catalog of large language model security research.

Update papers CI Explore License: MIT Data: JSON + CSV

Explore the web index | Browse all papers | Use the dataset | Methodology | Suggest a paper

SecPapers tracks both sides of LLM security: research that makes language models safer, and research that applies language models to cybersecurity. It queries arXiv every day, applies a transparent relevance filter, deduplicates paper revisions, and regenerates this repository from stable source data.

At a glance

1656 papers across 4 publication years. Latest arXiv metadata update: 2026-10-05.

Topic Papers
Prompt Injection & Jailbreaks 379
Agent & Tool Security 413
Privacy & Data Leakage 255
Safety, Alignment & Misuse 393
Adversarial ML, Poisoning & Backdoors 372
Software & Vulnerability Security 513
Malware, Phishing & Cyber Defense 204
Evaluation, Benchmarks & Red Teaming 798
Other LLM Security 105

Latest papers

Updated Paper Topics Links
2026-10-05 TranScope: What the Software Hides About LLM Training Data, the Hardware Reveals at Scale, and Accelerators Magnify
Joshua Kalyanapu, Darsh Asher, Kaushal Mhapsekar, et al.
Privacy & Data Leakage, Adversarial ML, Poisoning & Backdoors abstract / PDF
2026-10-05 Reward Stealing Attack on Large Language Models
Jiaming Qian, Pengyang Zhou, Jiahe Xu, et al.
Safety, Alignment & Misuse, Adversarial ML, Poisoning & Backdoors, Software & Vulnerability Security abstract / PDF
2026-10-05 Does AI Help Cyber Attackers or Defenders? Evidence from Nonpublic Vulnerabilities and Subsequent Attacks
Tobias Heldt, Matt Turk, Christoph Landolt, et al.
Software & Vulnerability Security abstract / PDF
2026-10-05 An Evaluation of the Semantic Understanding Capabilities of Large Language Models for Web Attack Payloads
Hao Sun, Yibin Yao, Chaohai Xie, et al.
Evaluation, Benchmarks & Red Teaming abstract / PDF
2026-10-05 Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Yongxiang Li, Moxin Li, Zhixin Ma, et al.
Agent & Tool Security, Evaluation, Benchmarks & Red Teaming abstract / PDF
2026-10-05 RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Mohamed Dhouib, Clement Elliker, Alexi Canesse, et al.
Prompt Injection & Jailbreaks, Agent & Tool Security, Adversarial ML, Poisoning & Backdoors abstract / PDF
2026-10-05 Correct Verdicts, Flawed Reasoning: Structured Auditing of LLM-based Vulnerability Reasoning
Boyue Caroline Hu, Kaivalya Ahir, Ronghao Ni, et al.
Software & Vulnerability Security, Evaluation, Benchmarks & Red Teaming abstract / PDF
2026-10-05 DP-ES: Differentially Private Evolution Strategies for Prompt Optimization
Ziniu Liu, Aiping Li, Yue Han, et al.
Privacy & Data Leakage, Adversarial ML, Poisoning & Backdoors, Evaluation, Benchmarks & Red Teaming abstract / PDF
2026-10-05 Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection
Takashi Koide, Hiroki Nakano, Daiki Chiba
Prompt Injection & Jailbreaks, Malware, Phishing & Cyber Defense, Evaluation, Benchmarks & Red Teaming abstract / PDF
2026-10-05 Where Did the Repair First Go Wrong? Localizing the Origins of Silent Failures in Agentic Vulnerability Repair
Wenji Bai, Muhammad Waseem, Zeeshan Rasheed, et al.
Agent & Tool Security, Software & Vulnerability Security, Evaluation, Benchmarks & Red Teaming abstract / PDF
2026-10-05 Benchmarking Jailbreak Guardrails for Embodied Agents
Xunguang Wang, Qingyue Wang, Yuguang Zhou, et al.
Prompt Injection & Jailbreaks, Safety, Alignment & Misuse, Evaluation, Benchmarks & Red Teaming abstract / PDF
2026-10-05 Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, et al.
Agent & Tool Security, Evaluation, Benchmarks & Red Teaming abstract / PDF
2026-10-05 Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PII
Alexandru Nazare, Agnese Profico, Nicolò Vania, et al.
Privacy & Data Leakage, Safety, Alignment & Misuse, Adversarial ML, Poisoning & Backdoors, Software & Vulnerability Security abstract / PDF
2026-10-05 Backdooring Sparse Autoencoders
Enrico Ahlers, Daniel Passon, Tobias Kiecker, et al.
Adversarial ML, Poisoning & Backdoors, Software & Vulnerability Security abstract / PDF
2026-10-05 PPFedIT: Towards Privacy-Preserving Federated Instruction Tuning with Few-shot Local Examples
Zhuo Zhang, Jingyuan Zhang, Jintao Huang, et al.
Privacy & Data Leakage, Software & Vulnerability Security abstract / PDF

Scope

Included work must mention an LLM or language-model concept and a concrete security, safety, privacy, abuse, or cyber-defense concept in its title or abstract. The taxonomy covers:

  • Prompt injection and jailbreaks
  • Agent and tool security
  • Privacy, memorization, and data leakage
  • Model safety, alignment, and misuse
  • Adversarial attacks, poisoning, and backdoors
  • Vulnerability discovery and secure software
  • Malware, phishing, and threat intelligence
  • Security evaluation, benchmarks, and red teaming

The catalog is automated discovery, not a quality ranking or endorsement. See the methodology for the query, scoring rules, known limitations, and correction process.

How it works

arXiv Atom API
      |
      v
query + pagination -> relevance scoring -> revision deduplication
      |                                          |
      +-------------------> data/papers.json <---+
                                  |
                                  v
            README.md + papers.md + CSV + web index

The collector uses only the Python standard library. There is no package installation step and no runtime dependency lockfile to maintain.

# Run tests
python3 -m unittest discover -s tests -v

# Fetch recent papers and regenerate every output
python3 scripts/collect.py

# Regenerate Markdown and CSV without network access
python3 scripts/collect.py --render-only

Search terms and taxonomy rules live in config/topics.json. The canonical record format is documented by data/schema.json. Updates run daily at 06:17 UTC and can also be started manually from the Actions tab.

Data use

  • data/papers.json is the canonical, stable dataset.
  • data/papers.csv is convenient for spreadsheets and analysis.
  • papers.md is the human-readable catalog grouped by topic.
  • docs/data contains compact, generated payloads for the SecPapers web index.
  • Each record links to the authoritative arXiv abstract and PDF.
  • Paper titles, abstracts, and author metadata remain attributable to their respective authors and are not relicensed by this repository's MIT license.

Contributing

False positives, missing papers, taxonomy improvements, and collector fixes are welcome. Read CONTRIBUTING.md before opening a pull request.

Acknowledgments

Paper metadata is provided by the arXiv API. SecPapers is not affiliated with or endorsed by arXiv. Please cite the original authors and papers when using this catalog in research.

About

A living catalog of LLM security research, updated daily from arXiv

Topics

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages