METER.

Benchmarks

These results come directly from METER's versioned evaluation record. Every reported result traces to a named, machine-readable run artifact.

Successful and failed evaluations are reported together. The model and evaluation criteria were prospectively locked before the evaluation data were examined. In each benchmark METER scores with no target-dataset fitting, while conventional reference models are fitted separately to the evaluation data where indicated.

METER is not intended to outperform a well-behaved raw total on every short scale, and on some of these instruments it does not. The runtime advantage is measured against fitted confirmatory psychometric models, which are re-estimated for every dataset; it is not a comparison with a sum score, which is effectively instantaneous.

Known truth

Unseen synthetic measurement worlds

METER r = 0.939 with known latent truth

Fitted GRM r = 0.944

Near-specialist recovery without target-specific fitting.

External transfer

66,812 respondents · 28 countries · EURO-D

METER vs fitted specialist r = 0.985

Weights locked before the evaluation

Scored without refitting, on a previously unseen real measurement problem.

Multidimensional

20,000 respondents · 50 items · 5 supplied factors

METER 0.78 s · confirmatory MIRT 662.1 s

Median agreement r = 0.979

Roughly 850× lower recorded runtime for pretrained scoring in this evaluation.

See the complete evidence, including failed evaluations ↓

EvaluationInstrumentNMETER vs referenceMETER fittingRuntime (reference vs METER)Evaluation outcome
Synthetic 1-D recovery, 40 unseen worldssynthetic GRM worlds (module M-ref)17,285person r 0.939 vs generating trait (mirt 0.944, raw sum 0.923)None3.9 s/world fitted vs 0.011 s/world pretrainedPASS
NHANES 2017–18, locked zero-shotPHQ-95,5330.9371 vs MML-GRM EAPNone9.7 s vs 0.12 ssupported
ESS Round 7, 21 countriesCES-D-840,185pooled 0.9696; per-country 0.9545–0.9899None11.2 s vs 0.56 ssupported
SHARE wave 9, locked, 28 countries (pre-registered gate)EURO-D66,8120.9851 [0.9849–0.9854]; 0.967–0.990 per countryNone13.8 s vs 6.4 sPASS
SHARE wave 8, locked temporal replicationEURO-D51,7320.9847; 27/27 countries 0.961–0.990None7.4 s vs 2.5 sPASS
IPIP Big-Five markers, locked, K = 5IPIP-FFM (module M-struct)20,000median 0.9791; per-factor 0.960–0.984 vs confirmatory MIRTNone662.1 s vs 0.78 sPASS
ICAR-16 cognition transfer, lockedICAR-164,574median 0.9023 vs the 0.93 gateNoneFAIL (permanent)
Longitudinal prospective noninferiority vs raw-sum mean (ELSA; SHARE-B), lockedCES-D-8; EURO-D10,805; 20,132deficit lower bounds −0.0237 / −0.0236 vs −0.02 marginNoneFAIL ×2 (preserved boundary)
Reliability indicator: conditional coverage (P5, truth-referenced)400 held-out synthetic worlds; 12 prespecified design strata400 worldscoverage 0.785–0.983; 10 of 12 strata outside the [0.93, 0.97] bandFailed the prespecified band with direction differing by response format (binary too narrow; 5- and 7-category too wide), so one global scalar calibration is not supported across those strata; the uncertainty output is therefore released as a relative reliability indicator (ordinal use only).NoneFAIL

Footnotes (integral to the table)

  1. Specialist comparators require dataset-specific optimization; METER performs one model pass with no per-dataset fitting.
  2. Raw sums remain strong for some simple instruments (ESS7 sum-agreement 0.998; the longitudinal raw-sum mean beat the prespecified noninferiority margin twice); METER is NOT claimed to universally outperform raw totals.
  3. Speed-ups labelled "recorded" are explicit artifact fields; "derived" are quotients of recorded runtimes; the largest ratios in the repository (D9C ~12,000x, D9B ~5,000x) come from NOT_EVALUABLE and FAIL runs and are deliberately excluded from this headline table.
  4. The D9C and D9E artifacts both record comparator fit 662.1 s (corroborated by both console logs); the ~850x claim additionally survives on the D9E numbers alone.

How to read these results

Agreement is not accuracy. On real data, agreement with a conventional estimator does not establish recovery of the true latent trait. METER and its comparator may share modelling assumptions and may therefore be wrong in similar ways. Only simulations with known generating values allow direct evaluation against latent truth. Measured: agreement of 0.990 can coexist with truth-recovery of 0.79–0.83.

Replication. The SHARE wave 8 result is a temporal replication within an overlapping longitudinal panel of 38,071 shared persons. It is not an independent-cohort replication.

View technical provenance

Evaluation record. Every number on this page is read from the versioned evidence matrix at repository tag METER_FOUNDATIONAL_EVIDENCE_FREEZE. Rows are never edited on this page; corrections happen upstream in the versioned record.

Outcome vocabulary. PASS denotes a pre-registered gate that was met, and FAIL a pre-registered gate that was not. Rows marked supported record an evaluation consistent with the capability where no pre-registered gate was set. These are evaluation-study outcomes and are separate from the capability decision a user receives for their own run.

Module attribution. METER is a release of separately trained, separately hash-pinned modules. The synthetic recovery row was produced by the 1-D reference core (M-ref). All real-data unidimensional rows come from the scoring core (M-score, 42,899 parameters, state 224a701b…). The IPIP row comes from the known-structure head (M-struct, state 8d427520…). The ICAR-16 failure is artifact D9D and the IPIP pass is artifact D9E.

Full methodological notes, per-row artifacts and the complete evidence base ship with the preprint (in preparation) and the model / capability card.

Run the 60-second demo