# Codebook v2 — evidence tiers (sharpened) + use-type

> **Status:** v2.1, adopted 2026-07-10 (Nick approved the sharpened boundary, the use-type
> dimension, and the consistency pass on 2026-07-08/10). v1 lives in `coding/codebook.md`
> and remains the rubric under which the original 2-coder census + gemini tiebreak ran;
> v1 labels are never edited. v2 is an enrichment layer coded separately
> (`processed/enrichment/`). Version history: v2.0 = sharpened tiers + use-type (sonnet
> enrichment pass, all 2,114 filings); v2.1 = v2.0 plus the "specific function counts"
> clause under tier 4 (GPT-5.5 boundary resolution + gemini tiebreak).

## What changed from v1, and why

1. **Tier 2 now explicitly includes "general own-use"** — a bare present-tense claim
   ("we use AI in our business", "we increasingly rely on AI") with no named system,
   workflow, or capability. Under v1 these scattered between tiers 1, 2, and 4
   (the instability Nick's 67-filing human review surfaced: Aligos, Customers Bancorp).
2. **Tier 4 requires a SPECIFIC capability.** v2.1 operationalizes "specific": a named
   system OR a specific described function ("AI-based tools that detect cybersecurity
   threats", "ML models that price small-commercial policies") counts; "we use AI to
   drive efficiencies" does not.
3. **New second dimension: USE-TYPE** — what the filing's strongest AI claim relates to.
   Separates companies that USE AI from companies that SELL AI (the Quarta-Rad case).

**Definitional decision (Nick, 2026-07-10):** a company that says it uses AI but cannot
name a single system or function gets tier 2, not production credit. The conservative
anti-hype stance: an unnameable capability is talk; `use_type` still records that the
filer claims own use.

## The tiers (v2.1)

| Tier | Name | Definition |
|---|---|---|
| 0 | Incidental / name-only | AI only in a company/product name, a bio, an auditor consent, or describing a third party, with no claim about the filer's own use. |
| 1 | Boilerplate / risk factor | Generic risk or competitive language (AI as threat, regulatory exposure, cyber risk). Item 1A defaults here unless it describes the filer's own deployed system. |
| 2 | Aspiration OR general own-use | Intends to / exploring / investing in AI, **or** a general present-tense own-use claim with **no specifics**. |
| 3 | Named pilot / initiative | A specific named tool, program, partnership, or pilot, not (yet) claimed in full production. |
| 4 | In production (specific) | A **specific** AI capability operating today — a named system **or** a specific described function — without a measured outcome. Bare "we use AI" is tier 2. |
| 5 | Quantified result | Tier 4 plus a specific measured outcome attributed to it. |

Filing-level unit: report the MAX tier across passages. Adjacent-tier ties break HIGH
(anti-bias: the project's thesis predicts low tiers).

## Use-type (primary classification of the strongest claim)

| Value | Meaning |
|---|---|
| `own_operations` | AI in the filer's own internal operations/processes. |
| `product_offering` | AI embedded in / sold as the filer's product, service, or platform. |
| `third_party` | AI as an external factor: competitor, vendor, regulation, market, threat. |
| `mixed` | Clearly both own_operations and product_offering. |

## The layered coding design (provenance)

1. **v1 census** (`coding/`): 2 blind coders (Claude Sonnet 5 + GPT-5.5), gemini
   tiebreak on disagreements, adjudicated labels in `coding/adjudication.jsonl`.
   Cohen's kappa 0.79, 86% agreement.
2. **Human review** (Nick, 67 stratified filings, review tool): 94% confirm rate.
3. **v2 enrichment** (`processed/enrichment/enrichment.jsonl`): one Sonnet coder,
   blind, all 2,114 filings, v2.0 rubric + use-type. Doubles as a consistency pass
   (89% same-bucket vs v1).
4. **Boundary resolution** (`resolution.jsonl`): GPT-5.5 blind under v2.1, only the
   203 filings where v1 and the enrichment disagree about production status.
5. **Tiebreak** (`tiebreak-v2.jsonl`): gemini blind under v2.1, only where sonnet and
   GPT-5.5 still disagree on the binary; majority of three settles it.
6. **Human re-review**: Nick re-reviews ~30 re-adjudicated boundary cases in the
   review tool before the v2 numbers publish.

Published v2 tier per filing: enrichment tier, overridden by the resolution/tiebreak
outcome for the contested set. v1 and v2 are both published columns in the open
dataset; the report states both definitions and their numbers.
