NLP models powering the OSINT platform at 667 posts/second: FastText lid.176 language detection (99.7% EN accuracy), custom SpaCy NER fine-tuned on 2.3M labeled examples across 7 political entity types (91.4% macro F1), DistilBERT fine-tuned on 5M examples with INT8 ONNX quantization (94.7% macro F1, 28ms GPU), MinHash character 4-gram coordinated-campaign detection (89% precision), and the social signal integration with Voidly censorship event detection.
September 2024
7 dated technical articles. These pages record the methods and claims as published; follow each article's source links and update notes for its current scope.
How the OSINT platform detects bot accounts across 14 languages without retraining per language: an 8-feature BotFeatureVector (posting_interval_entropy via Shannon formula, reply_outdegree_ratio, content_cluster_density, age_velocity_zscore, quote_to_original_ratio, url_recycling_rate, cross_platform_correlation, bio_change_count_90d), Redis-bucketed perceptual hash matching (Hamming ≤ 8 across 1024 hash buckets), XGBClassifier with StratifiedGroupKFold on language groups, and per-language Platt scaling achieving F1 0.883–0.908 across all 14 languages.
How we detect coordinated amplification campaigns across 58M daily posts: MinHash LSH (128 hash functions, 16 bands, Jaccard threshold 0.80) for content similarity, Redis sorted-set burst detection (≥5 accounts within 15 minutes, inverse-sqrt account age weighting), seven account-feature logistic regression, network amplification ring detection via cycle enumeration, cross-platform timing joins, and a 0–100 coordination score with 70/90 thresholds for human review and auto-flagging.
Machine learning and OSINT · Engineering and infrastructure · Money in politics
How the election intelligence pipeline resolves FEC committee identity across 1.3M records: the 10-code committee type taxonomy (H/S/P/X/Y/N/Q/O/I/U), a JointFundraisingCommittee dataclass with JFCAllocation and resolve_jfc_participants() from Form 99, normalize_entity_name() with iterative legal-suffix stripping, a four-pass resolution table (exact ID 63.4% → exact name 82.1% → alias 91.7% → TF-IDF char 3-gram 95.5% cumulative recall), and LLC chain disambiguation via FinCEN/EDGAR/SOS cross-reference.
Anomaly detection across 47 races in 23 states: Benford's law with magnitude-range validity checks, XGBoost turnout model (20 features, SHAP attribution, MAD-based z-scores, 3.1pp MAE), ARIMA(2,1,2) reporting-curve detection, DBSCAN campaign finance clustering (near-identical amounts + 3-day burst), and full triage workflow (12 flags → 9 explained, 2 false positives, 1 persistent).
The statistical methods behind AI Analytics' election anomaly detection — first-digit analysis, last-digit uniformity testing, turnout z-scores, and why these signals require cross-validation with social and media data before generating an alert.
Money in politics · Machine learning and OSINT · Censorship and information control
How the election intelligence pipeline ingests AP Election API feeds, state authority data (JSON/CSV/HTML scraping), social media signals, and media coverage in real time: Kafka election.precinct_results topic (50 partitions by state FIPS), PrecinctResult protobuf schema, state scraper StateScraperConfig, ElectionSentimentConsumer, narrative divergence scoring, FIPS normalization edge cases (Connecticut planning regions, Alaska districts), and p50/p99 latency targets for all four streams.
Money in politics · Engineering and infrastructure · Machine learning and OSINT