Evaluation & Testing
How we rigorously validate every model update before it reaches your clinic.
Evaluation Pipeline
Data Collection
Fresh clinical recordings collected monthly from partner clinics across 12 specialities and multiple Indian regions.
Gold Standard Annotation
Licensed physicians manually transcribe and structure notes. Two independent annotators per recording with inter-rater agreement >95%.
Automated Evaluation
AI outputs compared against gold standards using WER, clinical entity F1, section accuracy, and completeness metrics.
Clinical Review Board
Speciality physicians review edge cases, disagreements, and borderline outputs. Their judgements calibrate our automated metrics.
Regression Testing
Every model update runs through 10,000+ test cases before deployment. Any regression in any metric blocks the release.
Production Monitoring
Real-time quality signals from production (user edits, rejections, explicit feedback) feed back into evaluation priorities.
Evaluation Principles
Clinician Judges
Only licensed physicians evaluate clinical accuracy. never engineers alone.
Fresh Data Only
Test sets are never reused. Monthly refresh prevents overfitting to historical patterns.
Stratified Testing
Performance measured per-speciality, per-language, and per-accent to catch blind spots.
Zero-Regression Policy
No model ships if any metric degrades, even by 0.1%, without clinical review board approval.
Read the full methodology
Our whitepaper on clinical AI evaluation is available for download.
Download whitepaper