Full detail · part of The Platform

The Data Discipline

Anyone can accumulate data. The rare thing is being able to prove, row by row, where every number came from.

4.93M
medical encounters (defined)
6.67M
linked service items
99.999%
verified source traceability
2014–2026
observation period
Operational systems provider · pharmacy · lab every row journaled at source Journal + register source table registered, source key preserved Verification pass 4,934,458 of 4,934,493 events resolved to their source row 99.999% verified traceability Can every number be traced back to the row it came from? Yes — and it was machine-checked.
The lineage mechanism, machine-verified in June 2026: 4,934,458 of 4,934,493 events resolved to their originating source row; the 35 exceptions are individually documented.

How it works

The warehouse is journal-first: every operational system writes events through a journal that registers the source table and preserves the source key. That single design decision — made years before "data lineage" became a buzzword — is what makes verification possible: a script can walk every indexed event back to the exact row it came from and check that it exists. When I say 99.999% traceability, that's not a quality vibe; it's a measured, repeatable result.

Identity, honestly

Records about the same person arrive from provider, pharmacy, and laboratory systems with different identifiers. The identity layer anchors on strong keys where they exist and falls back to controlled matching where they don't — and, crucially, I report in units I can defend (encounters, membership episodes) rather than claiming unique-person counts the data can't prove. Refusing to publish a flattering-but-unprovable number is what governance means in practice.

Exceptions are catalogued, not hidden

Orphaned line items (0.41%), encounters flagged for treatment without billing detail (2.02%), date and demographic anomalies — all counted, classified, and excluded from any analysis they would distort. A dataset with documented flaws is trustworthy; a dataset with no reported flaws is unexamined.

Privacy by default

Pseudonymization at the identity layer, de-identification before any external use, aggregate-only exposure in anything web-facing, and tiered disclosure rules for sensitive figures. I completed U.S. medical billing & coding coursework in 2025 to map this practice onto U.S. vocabulary and privacy expectations.

Why this matters to you: if your organization has data that several systems disagree about — most do — this page is my job interview. The numbers above come from a validated profiling run I can walk you through in detail.