Research engineer with ~10 years of experience across LLM alignment research and production systems. I work on making models more honest and reliable — from preference optimization to RAG pipelines serving real users.
My research shifted to LLM honesty and calibration — models that give consistent answers and reflect their actual confidence
GRPO with custom reward signals to penalize hallucinated claims.
q-RAG: improving LLM coherence up to 28pp by augmenting prompts with equivalent questions. QUADRo: database QA over 6.3M pairs.
Billion-scale semantic similarity and ranking pipelines.
Multilingual data generation, large-scale annotation pipelines, quality filtering.
Preference optimization pipelines (DPO), human annotation systems for calibration evaluation, hallucination detection classifiers.
LLM calibration via RL (GRPO/DPO/ORPO, up to 72B params on 8×L40S), QA coherence, dataset curation. 6 publications.
Leading 10 engineers building Omnyscient, a multi-agent RAG platform. Hybrid retrieval, LangGraph, AWS Bedrock. Hundreds of thousands of users.
Scalable retrieval and ranking pipelines for database QA. Dense retrieval, billion-scale semantic similarity.
MCP plugin for ML experiment lifecycle: hypothesis-driven design, distributed training (multi-GPU, SSH), convergence detection, DPO/GRPO dataset validation. 478 tests.