Presentation
I design challenging clinical cases, evaluation rubrics, reference answers, and safety scenarios used to test AI systems. I evaluate model outputs for clinical accuracy, reasoning quality, evidence-based medicine, uncertainty, patient safety, contraindications, drug interactions, and guideline adherence.
My experience includes medical and multimodal AI evaluation, synthetic clinical case authoring, LLM red-teaming, agent trajectory analysis, clinical rubric design, medical-image evaluation, ECG interpretation, and quality assurance of healthcare AI training data.
I have worked on medical-AI evaluation projects through Outlier AI, Terac, Handshake AI, Alignerr and TELUS Digital, and I also operate an independent clinical evaluation practice focused on safety assessment and adversarial review of medical AI systems.
Available for: medical AI evaluation, LLM evaluation, clinical safety review, red-teaming, clinical dataset QA, medical benchmark creation, clinical scenario authoring, rubric design, and physician expert review.
Background
Design web-dependent questions intended to challenge frontier AI systems, identify reliable primary and secondary sources, and construct reference search trajectories demonstrating how a correct answer should be researched and verified.
Focus on evaluating both the quality of model answers and the reliability of the research process used to obtain them.
Construct end-to-end patient scenarios covering presentation, history, hidden clinical details, red flags, contraindications, drug interactions, allergies, comorbidities, and potential harm from incorrect AI guidance.
Author supporting synthetic medical documents including CBC, LFT, KFT, ABG, urine routine, toxicology reports, and prescriptions across clinical domains including autoimmune encephalitis, cardiology, and pulmonary embolism.
Each case undergoes individual clinical review and acceptance before payment, providing an external quality check for clinical accuracy.
Audit clinical rule sets and coded medical datasets against primary sources including ICD-10-CM Official Guidelines, FDA labeling, and current specialty guidelines, tracing clinical values to cited sources and documenting downstream consequences.
Conduct adversarial evaluation and responsible disclosure of medical AI failures, with a focus on clinical accuracy, patient safety, evidence-based reasoning, and model reliability.
Authored a research-proposal-style medical ML task defining the prediction target, explaining why domain expertise is important, and specifying train/test splits by parameter range to require genuine extrapolation.
Designed adversarial ML evaluation environments end to end, including problem domains, difficulty calibration against naive baselines, data schemas, and distribution-shift structures.
Directed AI-assisted implementation of underlying code and validated outputs against the intended task design, with accepted submissions.
Worked as a Maker on image-grounded visual question answering (VQA), authoring medical-domain prompts, reference responses, and structured explanations using an Introduce → Observe → Explain framework.
Evaluated generated content for factuality, uncertainty, and format adherence within a multi-stage Maker → Reviewer → Quality-Check pipeline and passed formal quality control on submitted work.
Worked up acute and undifferentiated clinical presentations in medicine, casualty, and critical care under senior supervision and interpreted ECGs on real patients, supported by formal ECG training from NPTEL, IIT Madras.
Rotated through radiology and laboratory medicine, interpreting imaging and laboratory results alongside clinicians and developing practical diagnostic-reasoning experience directly relevant to medical-imaging and diagnostic AI evaluation.
Designed healthcare-domain agent tasks and evaluation frameworks, including patient-vitals and medication-adherence agents. Built reference trajectories, authored scoring rubrics, and audited agent tool-call trajectories step by step.
Authored image-based medical prompts designed to stress-test frontier models and ranked multiple model responses across clinical and quality dimensions with written justification.
Designed rubric-based reward criteria for multi-step agent behavior, including RLVR tasks for tool-calling, defining both ideal responses and reasoning trajectories.
Iteratively rewrote prompts to expose model failures and created precise GUI click-target annotation tasks for computer-use AI evaluation.

