Benchmarks
These results come directly from METER's versioned evaluation record. Every reported result traces to a named, machine-readable run artifact.
Successful and failed evaluations are reported together. The model and evaluation criteria were prospectively locked before the evaluation data were examined. In each benchmark METER scores with no target-dataset fitting, while conventional reference models are fitted separately to the evaluation data where indicated.
METER is not intended to outperform a well-behaved raw total on every short scale, and on some of these instruments it does not. The runtime advantage is measured against fitted confirmatory psychometric models, which are re-estimated for every dataset; it is not a comparison with a sum score, which is effectively instantaneous.
Known truth
Unseen synthetic measurement worlds
METER r = 0.939 with known latent truth
Fitted GRM r = 0.944
Near-specialist recovery without target-specific fitting.
External transfer
66,812 respondents · 28 countries · EURO-D
METER vs fitted specialist r = 0.985
Weights locked before the evaluation
Scored without refitting, on a previously unseen real measurement problem.
Multidimensional
20,000 respondents · 50 items · 5 supplied factors
METER 0.78 s · confirmatory MIRT 662.1 s
Median agreement r = 0.979
Roughly 850× lower recorded runtime for pretrained scoring in this evaluation.
See the complete evidence, including failed evaluations ↓
| Evaluation | Instrument | N | METER vs reference | METER fitting | Runtime (reference vs METER) | Evaluation outcome |
|---|---|---|---|---|---|---|
| Synthetic 1-D recovery, 40 unseen worlds | synthetic GRM worlds (module M-ref) | 17,285 | person r 0.939 vs generating trait (mirt 0.944, raw sum 0.923) | None | 3.9 s/world fitted vs 0.011 s/world pretrained | PASS |
| NHANES 2017–18, locked zero-shot | PHQ-9 | 5,533 | 0.9371 vs MML-GRM EAP | None | 9.7 s vs 0.12 s | supported |
| ESS Round 7, 21 countries | CES-D-8 | 40,185 | pooled 0.9696; per-country 0.9545–0.9899 | None | 11.2 s vs 0.56 s | supported |
| SHARE wave 9, locked, 28 countries (pre-registered gate) | EURO-D | 66,812 | 0.9851 [0.9849–0.9854]; 0.967–0.990 per country | None | 13.8 s vs 6.4 s | PASS |
| SHARE wave 8, locked temporal replication | EURO-D | 51,732 | 0.9847; 27/27 countries 0.961–0.990 | None | 7.4 s vs 2.5 s | PASS |
| IPIP Big-Five markers, locked, K = 5 | IPIP-FFM (module M-struct) | 20,000 | median 0.9791; per-factor 0.960–0.984 vs confirmatory MIRT | None | 662.1 s vs 0.78 s | PASS |
| ICAR-16 cognition transfer, locked | ICAR-16 | 4,574 | median 0.9023 vs the 0.93 gate | None | — | FAIL (permanent) |
| Longitudinal prospective noninferiority vs raw-sum mean (ELSA; SHARE-B), locked | CES-D-8; EURO-D | 10,805; 20,132 | deficit lower bounds −0.0237 / −0.0236 vs −0.02 margin | None | — | FAIL ×2 (preserved boundary) |
| Reliability indicator: conditional coverage (P5, truth-referenced) | 400 held-out synthetic worlds; 12 prespecified design strata | 400 worlds | coverage 0.785–0.983; 10 of 12 strata outside the [0.93, 0.97] bandFailed the prespecified band with direction differing by response format (binary too narrow; 5- and 7-category too wide), so one global scalar calibration is not supported across those strata; the uncertainty output is therefore released as a relative reliability indicator (ordinal use only). | None | — | FAIL |
Footnotes (integral to the table)
- Specialist comparators require dataset-specific optimization; METER performs one model pass with no per-dataset fitting.
- Raw sums remain strong for some simple instruments (ESS7 sum-agreement 0.998; the longitudinal raw-sum mean beat the prespecified noninferiority margin twice); METER is NOT claimed to universally outperform raw totals.
- Speed-ups labelled "recorded" are explicit artifact fields; "derived" are quotients of recorded runtimes; the largest ratios in the repository (D9C ~12,000x, D9B ~5,000x) come from NOT_EVALUABLE and FAIL runs and are deliberately excluded from this headline table.
- The D9C and D9E artifacts both record comparator fit 662.1 s (corroborated by both console logs); the ~850x claim additionally survives on the D9E numbers alone.
How to read these results
Agreement is not accuracy. On real data, agreement with a conventional estimator does not establish recovery of the true latent trait. METER and its comparator may share modelling assumptions and may therefore be wrong in similar ways. Only simulations with known generating values allow direct evaluation against latent truth. Measured: agreement of 0.990 can coexist with truth-recovery of 0.79–0.83.
Replication. The SHARE wave 8 result is a temporal replication within an overlapping longitudinal panel of 38,071 shared persons. It is not an independent-cohort replication.
View technical provenance
Evaluation record. Every number on this page is read from the versioned evidence matrix at repository tag METER_FOUNDATIONAL_EVIDENCE_FREEZE. Rows are never edited on this page; corrections happen upstream in the versioned record.
Outcome vocabulary. PASS denotes a pre-registered gate that was met, and FAIL a pre-registered gate that was not. Rows marked supported record an evaluation consistent with the capability where no pre-registered gate was set. These are evaluation-study outcomes and are separate from the capability decision a user receives for their own run.
Module attribution. METER is a release of separately trained, separately hash-pinned modules. The synthetic recovery row was produced by the 1-D reference core (M-ref). All real-data unidimensional rows come from the scoring core (M-score, 42,899 parameters, state 224a701b…). The IPIP row comes from the known-structure head (M-struct, state 8d427520…). The ICAR-16 failure is artifact D9D and the IPIP pass is artifact D9E.
Full methodological notes, per-row artifacts and the complete evidence base ship with the preprint (in preparation) and the model / capability card.