7 dated technical articles. These pages record the methods and claims as published; follow each article's source links and update notes for its current scope.
How the Federal Regulatory Data Hub manages alias proliferation across OFAC SDN, SEC EDGAR, and FinCEN BSA: a five-type alias taxonomy (AKA/FKA/NFE/PHONETIC/VESSEL), entity_aliases DDL with FTS5 virtual table and covering indexes, a normalization pipeline with iterative legal-suffix stripping and NFKD ASCII transliteration, double-Metaphone phonetic bucket generation, and a four-pass resolution pipeline (exact 71.4% → phonetic 88.2% → FTS5 96.1% → edit-distance 98.7% cumulative recall on 2.4M aliases).
How the Federal Regulatory Data Hub resolves company identity across five incompatible federal identifier schemes: three-pass resolution strategy (exact ID join, alias table lookup, TF-IDF fuzzy name matching), the entity_master bridge table schema, company name normalization to remove legal suffixes, false positive rates by method, special cases for healthcare NPI arrays and foreign entities, and how the entity bridge achieves p50 38ms cross-agency query latency.
How the Federal Regulatory Data Hub implements full-text search across 50M+ records using SQLite FTS5 in Cloudflare D1: virtual table creation with the unicode61 tokenizer and content= shadow-table pattern, BM25 scoring with weighted columns (10× entity_name, 5× description, 1× narrative), highlight() and snippet() functions for context extraction, buildFts5Query() TypeScript alias expansion with legal suffix stripping, Promise.all cross-dataset fan-out across 5 D1 shards, trigger-based index maintenance, and weekly optimize via Cloudflare Cron.
How we built a 35M-record federal regulatory database on Cloudflare D1 — per-vertical SQLite tables across 208 datasets, daily cron ingest, FTS5 for free-text datasets, and vertical sharding past the 10GB limit.
How Voidly stores and queries 2.2 billion censorship probe results in TimescaleDB: hypertable design with 1-day chunk intervals and secondary country partitioning, 6.2× compression, continuous aggregates for country-level daily summaries, three-tier retention (hot/warm/cold), and query benchmarks for anomaly detection.