Skip to content

Blue Horizon Labs joins the Anthropic Claude Partner Network · Nord Security solutions partner

Blue Horizon Labs

Methodology · Digital Identity Index v1.0

The Digital Identity Index.

The Digital Identity Index measures identity integrity: whether a firm is provably the same, correctly-described, correctly-claimed entity everywhere its name appears. Eight dimensions, 100 points, quarterly, and zero judged dimensions — every check resolves to a verifiable fact or a typed evidence state.

The citable record

Instrument
Digital Identity Index (DII)
Version
v1.0
Dimensions
8
Maximum score
100
Cadence
Quarterly
Effective
2026-07-24
Lock window
2026-07-24 → 2026-11-20
Judged dimensions
0

These fields match the governing registry record exactly. The lock window is a commitment, not a plan: the rubric, the checklist, and the query battery were hash-frozen on the effective date and cannot change until the window closes. A reading taken in September and a reading taken in November are taken with the same ruler.

What it measures — and what it refuses to

One question: is this one entity, correctly stated and correctly owned, everywhere its name appears?

When a buyer, a partner portal, a knowledge graph, or a mail server encounters a firm’s name, do the facts, handles, marks, and machine-readable claims all resolve to one coherent, owned, correctly-stated entity — or does the identity fracture, drift, or collide with a homonym? That is the whole subject. It is a narrow question on purpose, because a narrow question is one you can answer without an opinion.

The index needs no external ground truth for consistency: the firm’s own primary domain is treated as canon, and every other surface is compared to it. Correctness is scored only against public records — the legal registry, DNS, and RDAP. No perception, no panel, and no aesthetic discretion enters the number.

It is also bounded away from its sibling instruments, deliberately and in writing. It scores whether a mark is present and the same everywhere, never whether the mark is good. It scores whether a handle is claimed and correctly attributed, never whether the account posts. It scores factual field parity, never tone or feel. It reads search results from human search engines with deterministic position counting, and never asks an AI answer engine how it would describe the firm — that belongs to the AI Visibility Index. It touches no performance, accessibility, or security grade, and no content or funnel outcome.

The rubric

Eight dimensions, 100 points, zero judged.

Class describes how the evidence is gathered: Auto from a machine query, Manual from a surface a person has to open, Hybrid from both. Every dimension is cross-firm — nothing in this rubric can only be scored on ourselves.

#DimensionWeightClassSource signal
01Canonical Facts & Contact Integrity18HybridName, email, phone, and address parity across every owned surface, against a frozen snapshot of the canonical facts.
02Namespace & Handle Integrity14HybridCore handles claimed and correctly owned — control, never activity — plus detection of confusable occupants nearby.
03Profile Completeness & Parity12ManualRequired identity fields present and consistent on each claimed profile.
04Visual Mark Presence & Consistency10HybridThe current mark present, and hash-matching the frozen reference, across owned surfaces.
05Machine-Readable Identity14AutoOrganization JSON-LD identity fields, sameAs ownership, and the identity block in llms.txt.
06Entity Disambiguation & Search Ownership12HybridWhether the owned entity owns its own name in human search results, counting homonyms that rank above it.
07Verification & Authenticity Signals12HybridPlatform verification and claim badges, SPF/DKIM/DMARC records, and RDAP/WHOIS coherence.
08Legal & Registry Identity Coherence8ManualLegal name, address, and status in the public business registry, against the canonical facts.
Total100Judged dimensions: 0.

The absence of a judged dimension is the design decision that matters most here. There is no pinned model, no calibration set, and no human scorer whose taste has to be controlled for — which means there is nothing in this instrument that could be tuned to flatter its author without changing the frozen rubric itself.

How a score is derived

The five-anchor ladder, applied per check.

Each checklist item inside a dimension carries an integer weight and resolves to exactly one anchor.

BandFractionMeaning at the check level
Exceptional1.00Check verified pass.
Strong0.80A graduated check, strong but not perfect — the tier is defined per item.
Adequate0.60A graduated check, partial.
Weak0.35Evidence could not be obtained, or a graduated check's weak tier.
Absent0.10Check verified fail. The floor is 0.10, never 0.00.

A dimension’s fraction is the weighted mean of its checks: the sum of each check’s anchor fraction times its weight, divided by the summed weight of every check that applied. Dimension points are that fraction times the dimension’s weight, and the total is the sum across all eight.

Because the failure anchor is 0.10 rather than 0.00, the effective floor of the whole index is 10, not 0. That is deliberate: a firm scoring 10 has been measured and found wanting on every check, and reporting that as “0” would read as “not measured.” The two are very different claims and the scale should not blur them.

If every check in a dimension turns out not to apply, the dimension is excluded entirely and its weight is redistributed proportionally across the survivors, preserving the hundred-point total exactly. Every such redistribution is flagged on the scorecard, and the applicable-weight denominator is always printed — a total that quietly changed its own basis is not comparable to anything.

Reported bands are read off the dimension fraction for presentation only and change no number: Exceptional at 0.90 and above, Strong from 0.70, Adequate from 0.50, Weak from 0.30, Absent below that. Each anchor sits cleanly inside its own band, so an all-pass dimension reports Exceptional and an all-fail dimension reports Absent, with no boundary ambiguity to argue about after the fact.

The evidence model

Four states. That is the whole apparatus.

StateFractionWhen it applies
pass1.00The canonical fact, claim, mark, or record was observed and it matches.
fail0.10The check was successfully performed and the observed value does not match — or a query that completed returned no such record.
not-applicableexcludedThe check does not apply to this firm. Its weight leaves the denominator; it is not scored as a failure.
unverifiable0.35The evidence could not be obtained — an auth-walled surface with no privileged access, or a transient tool or transport failure.

The fourth state is the one worth dwelling on, because it is where an instrument like this usually goes wrong. Any check whose evidence could not be obtained because of a tool or transport failure — a DNS timeout, an RDAP error, a network failure, a rate limit, a TLS handshake that never completed, an auth wall with no privileged access — resolves to unverifiable, never to a pass and never to a fail.

A negative finding may only be emitted from a query that succeededand returned no such record. A missing-DMARC finding produced by a DNS timeout is not a finding; it is a fabrication with a plausible shape, and it is the exact failure mode that makes automated scoring untrustworthy. Every unverifiable result carries its cause on the record, so a reader can tell the difference between “we looked and it was not there” and “we could not look.”

Pre-registration

What was frozen, before anything was scored.

The rubric above, the checklist that implements it, the platform set, the per-dimension check items, the pinned query battery for the search-ownership dimension, and the snapshot of canonical facts that everything is compared against were all hash-frozen on 2026-07-24, before a single Digital Identity Index score was recorded. With both the ruler and the definition of “correct” fixed in advance, the rubric cannot be tuned after the fact to flatter its author, and which facts count as canonical cannot be redefined once the score is known.

One disclosure has to be made plainly, and it is required on every artifact built on this instrument. The pre-registration is a rubric and threshold lock, nota blind prediction of an unknown outcome. Our own baseline defects were already surfaced during methodology design — the platform inventory that shaped the rubric is the same exercise that found them. So the integrity claim here is precisely this: we locked the ruler before we measured, and we publish whatever it reads. It is not, and must never be presented as, “we predicted a number sight unseen.”

The expectation recorded at the freeze — stated for honesty accounting rather than as a claim of foresight — was a baseline in the 55 to 70 range out of 100, with at least six typed findings. Deviation from that expectation publishes in either direction. A score above 70 and a score below 55 are both reported and explained; neither is quietly reconciled. Whether the cycle succeeded is not assessed here at all — that judgement is made in a second pre-registration frozen after the baseline locks, which separates findings we control from ceilings we do not.

What this does not prove

The reproducibility claim, stated at its real size.

For the automated and public-surface checks — DNS and email-authentication records, RDAP, Organization JSON-LD, the identity block in llms.txt, whether declared sameAs links resolve to surfaces the firm actually owns, mark bytes, and position counting in human search results — a third party holding the same frozen configuration can re-run the check and reproduce the result. Those checks are reproducible.

Version 1.0 makes no claim of being fully reproducible by a third party, and it would be false if it did. The manual and auth-walled checks — company-page and profile claim state, verification badges, field parity behind a login — are not symmetrically reproducible by an arbitrary outsider, because an outsider lacks privileged access to those surfaces.

That asymmetry cuts in a direction worth naming. A firm auditing itself can convert an auth-walled unverifiable check into a pass or a fail using privileged evidence; a firm being audited from outside cannot, and takes the weak-anchor credit instead. So a self-score on the auth-walled dimensions is potentially higher-confidence than an outside score on the same firm would be. Every self-assessed row is labelled as a self-assessment wherever it is published, for exactly that reason.

The first cycle

The first firm this was pointed at was us. The score is held pending review.

The rubric was promoted and made effective before any remediation work began, and that ordering is checkable against the commit history of the site it measures. The first cycle has been scored. It is not published yet, because publishing it means publishing a specific list of defects in our own identity surfaces, and that release is a decision the founder makes rather than one a publishing calendar makes for him.

No approximation stands in for it in the meantime. Not a rounded figure, not a band, not a direction of travel. Declining to publish a number is honest; publishing a comfortable version of it would end the usefulness of everything else on this page.

The benchmark hub carries the other three instruments and the rules the whole line runs under.

Begin

The same discipline, turned inward.

This instrument reads a firm’s identity from the outside. The diagnostic reads an operation from the inside — scored first, prescribed second, and published including the parts that are not flattering.

Schedule a diagnostic conversation