RESEARCH
Vijil Labs
You should not have to trust us
Vijil solves for trust in autonomous agents along four axes.
- Scale: the collective behavior of agent populations.
- Intervention: specification before development and verification after deployment, in a continuous loop.
- Metrics: alignment to specific principals rather than aggregated preferences.
- Mechanisms: changing the composition of agents as well as verifying their behavior.
TOPICS
What we work on
Adaptive red-teaming
Probes and attack strategies that surface hallucination, injection, jailbreak, and leakage failures — before adversaries do.
Trust measurement
Metrics and methodology for scoring reliability, security, and safety — with detectors calibrated against human-labeled ground truth.
Governance to guardrails
Small, fast open models for runtime defense — detecting prompt injections at inference time without adding latency.
Agent adaptation
Methods for attributing production failures to root causes and evolving agents under explicit constraints.
PUBLICATIONS & ARTIFACTS
Published work
FEATURED PAPEROpen Problems in Frontier AI Risk ManagementWhere the frontier risk-management process is still unresolved, and who is best placed to solve each gap.Ziosi, Plueckebaum, Casper, … Rudner (Vijil), … Trager · Oxford Martin AI Governance Initiative · Feb 2026Read the report →PAPERS
What the platform runs on
Each of these underwrites a mechanism you can go and check on a product page.
Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
nonfactuality detection
Lamb, Ivanova, Torr, Rudner (Vijil) · Preprint · arXiv:2604.07172 · 2026
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
nonfactuality detection
Bhardwaj, Kempe, Rudner (Vijil) · ICML 2026 · arXiv:2510.21891 · 2025
Red Teaming AI Red Teaming
the adversarial method
Majumdar (Vijil), Pendleton, Gupta · CAMLIS 2025 · arXiv:2507.05538 · 2025
Consistency in Language Models: Current Landscape, Challenges, and Future Directions
the reliability dimension
Novikova, Anderson, Blili-Hamelin, Majumdar (Vijil) · ICML 2025 workshop (R2-FM) · arXiv:2505.00268 · 2025
Improving Consistency in Large Language Models through Chain of Guidance
the reliability dimension
Raj, Gupta, Rosati, Majumdar (Vijil) · TMLR 2025 · arXiv:2502.15924 · 2025
Embedding-Based Classifiers Can Detect Prompt Injection Attacks
Dome’s runtime detectors
Ayub, Majumdar · CAMLIS 2024 · arXiv:2410.22284 · 2024
Is ETHICS About Ethics? Evaluating the ETHICS Benchmark
why a score is not evidence
Hancox-Li (Vijil), Blili-Hamelin · Preprint · arXiv:2410.13009 · 2024
Evaluating Defences Against Unsafe Feedback in RLHF
the safety dimension
Rosati, Edkins, Raj, Atanasov, Majumdar (Vijil), et al. · AAAI 2025 workshop (AICS) · arXiv:2409.12914 · 2024
garak: A Framework for Security Probing Large Language Models
Diamond’s probe framework
Derczynski, Galinkin, Martin, Majumdar (Vijil), Inie · Preprint · arXiv:2406.11036 · 2024
Representation Noising: A Defence Mechanism Against Harmful Finetuning
the safety dimension
Rosati, Wehner, Williams, et al., incl. Majumdar (Vijil) · NeurIPS 2024 · arXiv:2405.14577
Other work by our people
Published by Vijil researchers, and not part of the platform.
VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
Fang, Yuan, Kong, Rudner (Vijil) · ICLR 2026 · arXiv:2607.19575 · 2026
Open Problems in Frontier AI Risk Management
Ziosi, Plueckebaum, Casper, et al., incl. Rudner (Vijil) · Oxford Martin AI Governance Initiative · 2026
Localized LoRA: A Structured Low-Rank Approximation for Efficient Fine-Tuning
Barazandeh, Majumdar (Vijil), Rajyaguru, Michailidis · ICMLA 2025 · arXiv:2506.00236 · 2025
Stop Treating ‘AGI’ as the North-Star Goal of AI Research
Blili-Hamelin, Graziul, Hancox-Li (Vijil), et al. · ICML 2025 · arXiv:2502.03689
Unsocial Intelligence: An Investigation of the Assumptions of AGI Discourse
Blili-Hamelin, Hancox-Li (Vijil), Smart · AIES 2024 · arXiv:2401.13142 · 2024
Showing all 15 papers · sorted by recency
OPEN MODELS
Prompt-injection detection models
DeBERTa- and ModernBERT-based classifiers
The same detectors Dome runs, released for anyone to use, benchmark or fine-tune. They are the strongest thing on this page: you do not have to take our word about a detector you can download and point at your own traffic.
More research notes on the Vijil blog.
Try to beat our detector
Benchmark our prompt-injection detector against yours. It is on Hugging Face, it is small enough to run on your own hardware, and we would rather hear where it loses than where it wins — which is the only kind of marketing this page’s readers should accept.