One measurement layer. Many measures.
Turn supported questionnaire and assessment data into latent scores, diagnostics and an auditable analysis in seconds.
METER (Measurement Engine for Transferable Estimation and Representation) scores supported data with the same pretrained measurement model; it does not fit a new model to each dataset.
How METER works
Upload questionnaire data.
Choose items, coding and factor structure.
Compare METER with conventional scoring.
See how respondent rankings differ between methods.
Download scores, comparison report and Methods text.
Built for researchers who repeatedly measure latent characteristics
- Score a supported questionnaire without fitting another IRT model. Pretrained scoring replaces per-dataset estimation for supported designs.
- Run the same governed measurement workflow across different supported datasets and instruments. One Define · Check · Measure pipeline, one result object, one provenance record.
- Know when an analysis falls outside METER’s evidence rather than getting an unqualified number anyway. Limits and refusals are part of the product.
No per-dataset fitting
METER is trained once across simulated measurement problems and then left unchanged. For supported datasets, scoring is one model pass, with no fitting or recalibration on this dataset.
Refusals, not extrapolation
Before scoring, METER checks the request against a versioned capability contract, and the submitted responses against executable data checks: sample size, scale length, item variance, whether the items cohere, and per-factor adequacy. Outside the supported region METER limits or refuses rather than extrapolating.
Provenance on every result
Each run records the model, capability contract, data schema, item mapping, runtime and run ID, so every result can be audited backward to exactly what was asked and what answered.
Evaluated against conventional psychometrics
Known-truth recovery
On 40 unseen synthetic worlds with known latent truth, the pretrained core recovered person scores at median r = 0.939 (a per-dataset specialist reached 0.944).
Synthetic ground truth; real data has no truth to compare against.
External transfer without refitting
Prospectively locked before the data were opened: pooled agreement r = 0.985 with country-fitted models across 28 countries (66,812 respondents), replicated at 0.985 in an earlier wave.
Agreement with fitted models — convergence, not latent-truth accuracy.
Five-factor supplied structure
Factor-wise agreement 0.92–0.98 in an independent cohort, and median 0.979 when the same model weights scored a different 50-item instrument.
Scores transferred; factor-correlation recovery failed its prespecified gates, and METER never discovers structure.
See all benchmarks, including failed evaluations →
Starting with questionnaire and assessment data, METER is being built as a common measurement layer across supported instruments and populations. The current research release focuses on scoring, diagnostics, capability checks and auditable provenance.