Skip to content

Blue Horizon Labs joins the Anthropic Claude Partner Network · Nord Security solutions partner

Blue Horizon Labs

Open benchmarks

The rulers, published before the readings.

Blue Horizon Labs runs four scored instruments across identity, AI visibility, brand consistency, and marketing performance. This page is where their rubrics live — the dimensions, the weights, the scoring rules, and the lock window each one is frozen under. You can disagree with a measurement here without having to guess how it was taken.

Why the rubrics are public

A rubric is not a moat. Publishing it costs nothing and buys the one thing a new domain cannot otherwise buy.

The usual reason a firm keeps its scoring method private is that the method is the product. Ours is not. What a rubric buys you is the ability to check our work — to take the same eight dimensions and the same weights, run them against the same public evidence, and see whether you get the number we got. A measurement nobody can reproduce is an opinion with a decimal point.

That is also the honest answer to why this page exists at all. Blue Horizon Labs is the Capital Region’s applied-AI research lab, and a lab that publishes only its conclusions is a consultancy with a nicer word for itself. The instruments below are versioned, dated, and locked for a stated window precisely so that a reading taken in October can be compared against one taken in July without anyone having quietly moved the ruler in between.

As of 2026-07-25, we are the only Upstate New York AI firm publishing open benchmarks. That line is a finding rather than a boast, and it comes with an expiry attached: it is re-checked against the regional competitive set every month, and the moment another firm publishes a rubric it narrows — publicly, in the same cycle, on this page. A claim you would not let anyone falsify is not a claim; it is decoration.

The instruments

Four instruments, and what each one reads.

Each block below is the citable record: instrument, version, dimension count, maximum score, cadence, effective date, and the window the rubric is locked under. They match the governing registry records exactly. Where they ever disagree, the registry is right and the page is wrong.

Digital Identity Index (DII)

Measures identity integrity: whether a firm is provably the same, correctly-described, correctly-claimed entity everywhere its name appears.

Instrument
Digital Identity Index (DII)
Version
v1.0
Dimensions
8
Maximum score
100
Cadence
Quarterly
Effective
2026-07-24
Lock window
2026-07-24 → 2026-11-20
Judged dimensions
0

AI Visibility Index (AIVI)

Measures how well a firm survives the shift from human search to AI-mediated discovery.

Instrument
AI Visibility Index (AIVI)
Version
v1.0
Dimensions
8
Maximum score
100
Cadence
Monthly
Effective
2026-05-20
Lock window
2026-05-20 → 2026-11-20

Rubric page in preparation. Until it is up, this record is what we will stand behind: the version and the lock window are the same ones the instrument is running under today, not a plan.

Brand Consistency Index (BCI)

Measures the brand consistency and quality of a consulting firm as perceived by a sophisticated mid-market buyer.

Instrument
Brand Consistency Index (BCI)
Version
v1.1
Dimensions
8
Maximum score
100
Cadence
Quarterly
Effective
2026-05-20
Lock window
2026-05-20 → 2026-11-20

Rubric page in preparation. Until it is up, this record is what we will stand behind: the version and the lock window are the same ones the instrument is running under today, not a plan.

Marketing Performance Index (MPI)

Measures marketing engine performance — the operational signals that indicate whether a firm's marketing is producing pipeline, audience growth, or market authority now.

Instrument
Marketing Performance Index (MPI)
Version
v1.0
Dimensions
8
Maximum score
100
Cadence
Monthly
Effective
2026-05-20
Lock window
2026-05-20 → 2026-11-20

Internal instrument — two of its eight dimensions read our own pipeline and do not transfer across firms. It is listed here for completeness rather than hidden, because an instrument set with a quiet gap in it invites the question of what else is missing.

The self-audit

The first instrument we pointed at ourselves.

The Digital Identity Index was locked on 2026-07-24 and the first cycle it scored was our own. That ordering is deliberate and it is checkable: the rubric’s promotion date and its effective-from date both precede every remediation commit we have made since. Locking the ruler before measuring, and publishing whatever it reads, is the whole integrity claim — not that we predicted a number.

The first cycle’s score is held pending review. It has been taken, against the pre-registered expectation published in the rubric, and it is not on this page yet. Releasing it means releasing a specific list of defects in our own identity surfaces, and that release is a decision the founder makes rather than one a publishing schedule makes for him.

What will not happen in the meantime is a rounded, softened, or approximate figure standing in for the real one. Declining to publish a number is honest. Publishing a flattering version of it would be the exact failure this whole line exists to avoid, and it would be the last time anything on this page was worth citing.

The rules

What we hold ourselves to when we publish a measurement.

  1. 01

    Methodology before scores, always.

    Every rubric is published before, or at minimum alongside, any score derived from it. A published rubric with no score is honest work in progress. A score with no rubric is a marketing number.

  2. 02

    Version, effective date, and lock window on every page.

    Each matches the governing registry record exactly. If the page and the record disagree, the page is wrong — and the block at the top of each rubric page is where you check.

  3. 03

    No firm is scored publicly before its rubric is public.

    And no regional firm is named and shamed. Our regional study named competitors only on publicly observable activity, and left the tracked advisory firms that showed no AI moves unnamed. That precedent holds.

  4. 04

    No client appears on a benchmark page.

    Client work is described by industry rather than by name on every outward surface, and a rubric has no reason to reference an engagement at all.

  5. 05

    Deviations publish in both directions.

    A pre-registered expectation that is missed gets reported and explained, not quietly reconciled — and a result that beats the expectation is not treated as more newsworthy than one that misses it.

These sit on top of the lab’s wider research standard, which pre-registers a hypothesis, a baseline, a method, and a success threshold before a study runs, and then publishes the outcome whether it confirmed, partly confirmed, came back null, contradicted the hypothesis, or is still open. Results from client engagements are anonymised, and a non-positive client outcome is published only with that client’s explicit consent — absent consent, the honest record stays internal and the study does not run in public at all.

What this is not

These pages convert nothing, and should not be judged as if they did.

A methodology page is not a lead-generation asset. Nobody reads a weight table and books a diagnostic. These exist to be cited, linked, and used as corroborating evidence when someone — a person or a retrieval system — is deciding whether “applied-AI research lab” is a description or a costume. Judged on traffic they will look like failures. Judged on whether the work can be checked, they are the only assets here that compound.

If you want the readings rather than the rulers, the regional study of AI across the Capital Region is the instrument turned outward, with its funnel counts and its exclusions published alongside the findings. The rest of the library is the arguments those readings produced.

Begin

Measured the same way, on your operation.

The instruments on this page read firms from the outside. The diagnostic reads a business from the inside — same discipline, same habit of publishing the parts that are not flattering.

Schedule a diagnostic conversation