Admin console

MVP

Metrics and evaluation

Read-only. Every figure below is an engineering target, not a measured result: nothing is built yet, and a target quoted as an outcome would be a lie a judge could catch.

No evaluation run has been scored yet. The gold set is drafted at 60 items. These are the thresholds the system is being built to meet, stated in advance so they can be checked against it later.

MeasureTargetWhy it is the target
Citation precisionTarget: 100 percent of rendered citations resolve to a verified spanA single fabricated citation fails the product's core claim, so the target is absolute rather than high.
Answer accuracy on the gold setTarget: 85 percent on a 60-item gold setGold set drafted, not yet scored.
Correct abstention rateTarget: 95 percent of out-of-scope queries abstain rather than answerMeasured against deliberately unanswerable probes.
Over-abstention rateTarget: below 10 percentThe failure mode on the other side: refusing questions it could have answered.
Wrong-with-confidence rateTarget: 0 answers rendered Verified that are materially wrongThe single most damaging failure mode for a trust product.
Cross-lingual agreementTarget: 0.9 Jaccard between English and Hindi renderingsGate 8 blocks delivery below threshold.
Latency, medianTarget: under 8 seconds end to end on CPU-only hardwareGeneration dominates. Hardware acceleration is measured separately.

Pramaan gate outcomes, most recent answer

10 of 11 gates passed. One is not applicable because the answer was served in English.

  • Gate 1Span existencepass
  • Gate 2Offset integritypass
  • Gate 3Version currencypass
  • Gate 4Jurisdiction puritypass
  • Gate 5Authority tierpass
  • Gate 6Entailmentpass
  • Gate 7Temporal orderingpass
  • Gate 8Cross-lingual checkN/A
  • Gate 9Compositional ledgerpass
  • Gate 10Primary versus commentarypass
  • Gate 11Image-bound citationpass