New: Voice AI now handles 47+ calls daily per clinic. See how it works

Evaluation & Testing

How we rigorously validate every model update before it reaches your clinic.

Evaluation Pipeline

1

Data Collection

Fresh clinical recordings collected monthly from partner clinics across 12 specialities and multiple Indian regions.

2

Gold Standard Annotation

Licensed physicians manually transcribe and structure notes. Two independent annotators per recording with inter-rater agreement >95%.

3

Automated Evaluation

AI outputs compared against gold standards using WER, clinical entity F1, section accuracy, and completeness metrics.

4

Clinical Review Board

Speciality physicians review edge cases, disagreements, and borderline outputs. Their judgements calibrate our automated metrics.

5

Regression Testing

Every model update runs through 10,000+ test cases before deployment. Any regression in any metric blocks the release.

6

Production Monitoring

Real-time quality signals from production (user edits, rejections, explicit feedback) feed back into evaluation priorities.

Evaluation Principles

Clinician Judges

Only licensed physicians evaluate clinical accuracy. never engineers alone.

Fresh Data Only

Test sets are never reused. Monthly refresh prevents overfitting to historical patterns.

Stratified Testing

Performance measured per-speciality, per-language, and per-accent to catch blind spots.

Zero-Regression Policy

No model ships if any metric degrades, even by 0.1%, without clinical review board approval.

Read the full methodology

Our whitepaper on clinical AI evaluation is available for download.

Download whitepaper
Chat with us on WhatsApp