Dylan Bouchard, PhD

Lead Applied Research Scientist, Thomson Reuters Labs

I'm an applied research scientist at Thomson Reuters Labs. My research focuses on AI safety, particularly uncertainty quantification, hallucination detection, and bias and fairness in large language models. Previously, I built and led the AI Research program at CVS Health, conducting research and developing open-source projects including UQLM and LangFair. I hold a PhD in Economics (Econometrics) from North Carolina State University.

Select Publications

  1. D. Bouchard, M. S. Chauhan, D. Skarbrevik, H.-K. Ra, V. Bajaj, and Z. Ahmad. UQLM: A Python Package for Uncertainty Quantification in Large Language Models. Journal of Machine Learning Research, 27(13):1–10, 2026. [JMLR]
  2. D. Bouchard, M. S. Chauhan, Z. Ahmad, and H.-K. Ra. Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification. EMNLP (Main Conference), 2026. [arXiv]
  3. D. Bouchard and M. S. Chauhan. Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers. Transactions on Machine Learning Research, 2025. [TMLR]
  4. D. Bouchard, M. S. Chauhan, V. Bajaj, and D. Skarbrevik. Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Study. Transactions on Machine Learning Research, 2026. [TMLR]
  5. M. S. Chauhan, V. Gyanchandani, and D. Bouchard. When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study. Findings of AACL-IJCNLP, 2026. [arXiv]
  6. D. Bouchard. Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs. LT-EDI Workshop at ACL, 2026. [ACL Anthology]
  7. D. Bouchard, M. S. Chauhan, D. Skarbrevik, V. Bajaj, and Z. Ahmad. LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases. Journal of Open Source Software, 10(105):7570, 2025. [DOI]
  8. D. Bouchard and M. S. Chauhan. Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents. Under review, 2026. [arXiv]
  9. D. Bouchard. Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascades. Under review, 2026. Accepted to the AdaptFM Workshop at ICML 2026 (non-archival). [arXiv]

Open-Source Projects

Selected Talks

Service

Reviewing: NeurIPS, ICML, ICLR, COLM, ACL, ACM TIST, and ACM CHI.

Program committee: Multiplicity and Homogenization in AI (NeurIPS 2026).

Awards: Best Reviewer, LT-EDI Workshop at ACL 2026.