methodology // readability evidence
How we measure the lift.
A real-user visibility headline appears only when a fresh, production-parity benchmark passes every publication gate. Separately labelled synthetic retrieval diagnostics may be shared as benchmark-specific research. This page explains what each result measures and what it does not claim.
01 // the experiment
Matched real-user A/B. Each participant supplies an original PDF, DOCX, TXT or platform export and approves which source-present facts may be public. The baseline is the normal production extraction of that exact file. The optimized side is Humetric’s exposure-filtered SSR, ProfilePage/Person JSON-LD and public evidence chunks. Neither side may delete source facts or receive evaluation-only facts.
Human-reviewed gold facts. Two blinded reviewers adjudicate atomic identity, role, location, years, skills, experience, education and language facts. A fact enters the primary score only when it is present in the source and approved for public exposure. Hidden and contact facts are excluded from the score and covered by a zero-tolerance privacy test.
Primary metric. Machine-readable Public Fact Recall (MPFR) is the share of approved gold facts recovered by the same frozen extractor from each representation. Relative lift is 100 × (MPFRHumetric − MPFRoriginal) / MPFRoriginal. A zero baseline has no relative percentage; only the percentage-point change may be reported.
Uncertainty and release gate. The 95% interval comes from 10,000 seeded paired bootstrap iterations over people. Publication requires at least 100 consented profiles, PDF and DOCX strata of at least 30 each, at least 30 Traditional-Chinese/bilingual profiles, 40%+ primary lift, optimized MPFR ≥ 0.80, a positive CI, source-fact parity ≥ 0.99, zero privacy/critical Schema failures and no material subgroup regression.
02 // public retrieval diagnostics
| comparison | baseline → optimized | result (95% CI) |
|---|---|---|
| Fair same-fact A/Bretrieval quality · synthetic | 0.804 → 0.828 | +3.0% (+0.4% to +5.8%) |
| Documented synthetic prototypecomposite retrieval | 0.802 → 0.904 | +12.7% (+9.6% to +16.2%) |
| Documented synthetic prototypebounded ranking loss (1 − R) | 0.198 → 0.096 | 51.4% lower ( 43.2% to 59.3%) |
All three figures use the same 300-profile synthetic corpus and 155 frozen top-10 queries. The fair A/B changes representation only while holding source facts and the retrieval stack constant. The prototype compares a same-fact single document with field-aware chunks and typed-facet reranking. Benchmark-specific retrieval results—not a guarantee of crawling, indexing, ranking, citation, or visibility on any specific AI or search platform. Reported prototype intervals are documented research results; reproducible raw run artifacts are pending archival.
03 // current real-user result
04 // what this does NOT claim
The pre-registered specification lives at AI_VISIBILITY_BENCHMARK_V1.md. Synthetic diagnostics may be shared only with explicit synthetic and benchmark-specific labels; they cannot populate the real-user MPFR visibility claim gate. The +12.7% and 51.4% prototype intervals remain documented results until their reproducible raw run artifacts are archived.