
Divyanshu Mishra
Delhi, India
Divyanshu Mishra
MBBS Physician | Medical AI Evaluation
Category : Doctor
I am an MBBS physician specializing in medical AI evaluation, clinical reasoning assessment, adversarial red-teaming, and clinical safety review of healthcare AI systems and LLMs.
I design challenging clinical cases, evaluation rubrics, reference answers, and safety scenarios used to test AI systems. I evaluate model outputs for clinical accuracy, reasoning quality, evidence-based medicine, uncertainty, patient safety, contraindications, drug interactions, and guideline adherence.
My experience includes medical and multimodal AI evaluation, synthetic clinical case authoring, LLM red-teaming, agent trajectory analysis, clinical rubric design, medical-image evaluation, ECG interpretation, and quality assurance of healthcare AI training data.
I have worked on medical-AI evaluation projects through Outlier AI, Terac, Handshake AI, Alignerr and TELUS Digital, and I also operate an independent clinical evaluation practice focused on safety assessment and adversarial review of medical AI systems.
Available for: medical AI evaluation, LLM evaluation, clinical safety review, red-teaming, clinical dataset QA, medical benchmark creation, clinical scenario authoring, rubric design, and physician expert review.
I design challenging clinical cases, evaluation rubrics, reference answers, and safety scenarios used to test AI systems. I evaluate model outputs for clinical accuracy, reasoning quality, evidence-based medicine, uncertainty, patient safety, contraindications, drug interactions, and guideline adherence.
My experience includes medical and multimodal AI evaluation, synthetic clinical case authoring, LLM red-teaming, agent trajectory analysis, clinical rubric design, medical-image evaluation, ECG interpretation, and quality assurance of healthcare AI training data.
I have worked on medical-AI evaluation projects through Outlier AI, Terac, Handshake AI, Alignerr and TELUS Digital, and I also operate an independent clinical evaluation practice focused on safety assessment and adversarial review of medical AI systems.
Available for: medical AI evaluation, LLM evaluation, clinical safety review, red-teaming, clinical dataset QA, medical benchmark creation, clinical scenario authoring, rubric design, and physician expert review.
Working hours
- Monday:08h00 To 18h00
- Tuesday:08h00 To 18h00
- Wednesday:08h00 To 18h00
- Thursday:08h00 To 18h00
- Friday:08h00 To 18h00
- Saturday:Not available
- Sunday:Not available
Selected following assessment in both attempter and reviewer tracks for web-grounded frontier-model evaluation.
Design web-dependent questions intended to challenge frontier AI systems, identify reliable primary and secondary sources, and construct reference search trajectories demonstrating how a correct answer should be researched and verified.
Focus on evaluating both the quality of model answers and the reliability of the research process used to obtain them.
Design web-dependent questions intended to challenge frontier AI systems, identify reliable primary and secondary sources, and construct reference search trajectories demonstrating how a correct answer should be researched and verified.
Focus on evaluating both the quality of model answers and the reliability of the research process used to obtain them.
Author synthetic clinical cases used as training and evaluation data for medical AI systems, with each case designed around realistic patients, clinical conditions, and safety constraints.
Construct end-to-end patient scenarios covering presentation, history, hidden clinical details, red flags, contraindications, drug interactions, allergies, comorbidities, and potential harm from incorrect AI guidance.
Author supporting synthetic medical documents including CBC, LFT, KFT, ABG, urine routine, toxicology reports, and prescriptions across clinical domains including autoimmune encephalitis, cardiology, and pulmonary embolism.
Each case undergoes individual clinical review and acceptance before payment, providing an external quality check for clinical accuracy.
Construct end-to-end patient scenarios covering presentation, history, hidden clinical details, red flags, contraindications, drug interactions, allergies, comorbidities, and potential harm from incorrect AI guidance.
Author supporting synthetic medical documents including CBC, LFT, KFT, ABG, urine routine, toxicology reports, and prescriptions across clinical domains including autoimmune encephalitis, cardiology, and pulmonary embolism.
Each case undergoes individual clinical review and acceptance before payment, providing an external quality check for clinical accuracy.
Founder of an independent clinical evaluation practice focused on medical AI safety assessment, adversarial review, evaluation rubric design, and guideline-anchored clinical content authoring.
Audit clinical rule sets and coded medical datasets against primary sources including ICD-10-CM Official Guidelines, FDA labeling, and current specialty guidelines, tracing clinical values to cited sources and documenting downstream consequences.
Conduct adversarial evaluation and responsible disclosure of medical AI failures, with a focus on clinical accuracy, patient safety, evidence-based reasoning, and model reliability.
Audit clinical rule sets and coded medical datasets against primary sources including ICD-10-CM Official Guidelines, FDA labeling, and current specialty guidelines, tracing clinical values to cited sources and documenting downstream consequences.
Conduct adversarial evaluation and responsible disclosure of medical AI failures, with a focus on clinical accuracy, patient safety, evidence-based reasoning, and model reliability.
Contributed to frontier-model evaluation as a medical-domain expert, designing challenging problems where clinical and diagnostic knowledge is essential for solving the task.
Authored a research-proposal-style medical ML task defining the prediction target, explaining why domain expertise is important, and specifying train/test splits by parameter range to require genuine extrapolation.
Designed adversarial ML evaluation environments end to end, including problem domains, difficulty calibration against naive baselines, data schemas, and distribution-shift structures.
Directed AI-assisted implementation of underlying code and validated outputs against the intended task design, with accepted submissions.
Authored a research-proposal-style medical ML task defining the prediction target, explaining why domain expertise is important, and specifying train/test splits by parameter range to require genuine extrapolation.
Designed adversarial ML evaluation environments end to end, including problem domains, difficulty calibration against naive baselines, data schemas, and distribution-shift structures.
Directed AI-assisted implementation of underlying code and validated outputs against the intended task design, with accepted submissions.
Selected as a medical-domain contributor in Clinical Medicine and Basic Medical Science, authoring and quality-checking training data for frontier large language models with a focus on diagnostics and laboratory medicine.
Worked as a Maker on image-grounded visual question answering (VQA), authoring medical-domain prompts, reference responses, and structured explanations using an Introduce → Observe → Explain framework.
Evaluated generated content for factuality, uncertainty, and format adherence within a multi-stage Maker → Reviewer → Quality-Check pipeline and passed formal quality control on submitted work.
Worked as a Maker on image-grounded visual question answering (VQA), authoring medical-domain prompts, reference responses, and structured explanations using an Introduce → Observe → Explain framework.
Evaluated generated content for factuality, uncertainty, and format adherence within a multi-stage Maker → Reviewer → Quality-Check pipeline and passed formal quality control on submitted work.
Completed the 12-month Compulsory Rotatory Residential Internship required after MBBS, rotating through internal medicine, general surgery, casualty/emergency, anaesthesia and critical care, community medicine, obstetrics and gynaecology, paediatrics, radio-diagnosis, and laboratory medicine.
Worked up acute and undifferentiated clinical presentations in medicine, casualty, and critical care under senior supervision and interpreted ECGs on real patients, supported by formal ECG training from NPTEL, IIT Madras.
Rotated through radiology and laboratory medicine, interpreting imaging and laboratory results alongside clinicians and developing practical diagnostic-reasoning experience directly relevant to medical-imaging and diagnostic AI evaluation.
Worked up acute and undifferentiated clinical presentations in medicine, casualty, and critical care under senior supervision and interpreted ECGs on real patients, supported by formal ECG training from NPTEL, IIT Madras.
Rotated through radiology and laboratory medicine, interpreting imaging and laboratory results alongside clinicians and developing practical diagnostic-reasoning experience directly relevant to medical-imaging and diagnostic AI evaluation.
Oracle-tier contributor on medical and agentic AI evaluation projects, recognized for consistent high-quality work.
Designed healthcare-domain agent tasks and evaluation frameworks, including patient-vitals and medication-adherence agents. Built reference trajectories, authored scoring rubrics, and audited agent tool-call trajectories step by step.
Authored image-based medical prompts designed to stress-test frontier models and ranked multiple model responses across clinical and quality dimensions with written justification.
Designed rubric-based reward criteria for multi-step agent behavior, including RLVR tasks for tool-calling, defining both ideal responses and reasoning trajectories.
Iteratively rewrote prompts to expose model failures and created precise GUI click-target annotation tasks for computer-use AI evaluation.
Designed healthcare-domain agent tasks and evaluation frameworks, including patient-vitals and medication-adherence agents. Built reference trajectories, authored scoring rubrics, and audited agent tool-call trajectories step by step.
Authored image-based medical prompts designed to stress-test frontier models and ranked multiple model responses across clinical and quality dimensions with written justification.
Designed rubric-based reward criteria for multi-step agent behavior, including RLVR tasks for tool-calling, defining both ideal responses and reasoning trajectories.
Iteratively rewrote prompts to expose model failures and created precise GUI click-target annotation tasks for computer-use AI evaluation.
MBBS physician with clinical training across internal medicine, emergency/casualty, surgery, anaesthesia & critical care, paediatrics, OBG, radiology and laboratory medicine. Developed strong foundations in clinical reasoning, ECG interpretation, diagnostic assessment and evidence-based medicine. This clinical background now supports my work in medical AI evaluation, clinical safety assessment, adversarial red-teaming and evaluation rubric design for healthcare AI systems.
Please sign in as a customer to give your feedback


