Integra Lab
We build benchmarks and evaluations to find where language models fail.
Led by Ashwin Kirubakaran and Henry Gagnier
Publications
- BioConflict: A Benchmark for Evaluating Large Language Models in Biomedical Contradiction Detection and Consensus SynthesisBioNLP @ ACL 2026
- Deer, Deities, and Dancing: Culturally Biased LLM Hallucination in Low-Resource Wixárika TranslationAmericasNLP @ ACL 2026
- A Benchmark and Evaluation of Automated Language of Study Extraction from Computational Linguistics PublicationsSRW @ EACL 2026
- KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCRAbjadNLP @ EACL 2026