Admin console
MVPMetrics and evaluation
Read-only. Every figure below is an engineering target, not a measured result: nothing is built yet, and a target quoted as an outcome would be a lie a judge could catch.
| Measure | Target | Why it is the target |
|---|---|---|
| Citation precision | Target: 100 percent of rendered citations resolve to a verified span | A single fabricated citation fails the product's core claim, so the target is absolute rather than high. |
| Answer accuracy on the gold set | Target: 85 percent on a 60-item gold set | Gold set drafted, not yet scored. |
| Correct abstention rate | Target: 95 percent of out-of-scope queries abstain rather than answer | Measured against deliberately unanswerable probes. |
| Over-abstention rate | Target: below 10 percent | The failure mode on the other side: refusing questions it could have answered. |
| Wrong-with-confidence rate | Target: 0 answers rendered Verified that are materially wrong | The single most damaging failure mode for a trust product. |
| Cross-lingual agreement | Target: 0.9 Jaccard between English and Hindi renderings | Gate 8 blocks delivery below threshold. |
| Latency, median | Target: under 8 seconds end to end on CPU-only hardware | Generation dominates. Hardware acceleration is measured separately. |
Pramaan gate outcomes, most recent answer
10 of 11 gates passed. One is not applicable because the answer was served in English.
- Gate 1Span existencepass
- Gate 2Offset integritypass
- Gate 3Version currencypass
- Gate 4Jurisdiction puritypass
- Gate 5Authority tierpass
- Gate 6Entailmentpass
- Gate 7Temporal orderingpass
- Gate 8Cross-lingual checkN/A
- Gate 9Compositional ledgerpass
- Gate 10Primary versus commentarypass
- Gate 11Image-bound citationpass