Build & technical background — architecture decisions, design rationale, spec pointers. The context an implementation agent reads before planning work.
4 notes logged
+ Paste new session
Conflict pipeline audit 2026-08-08 + Thread 1/2 execution plan (repo: docs/CWP-Pilot-and-Report-Card-Execution-Plan.md)
2026-08-08
Edit

Full review of the conflict-pipeline build and forward plan. Canonical doc: docs/CWP-Pilot-and-Report-Card-Execution-Plan.md (branch docs-cwp-pilot-plan, pushed; also on disk in the working tree). Read that before touching pilot or report-card work.

PIPELINE STATE (DB-audited 2026-08-08):

  • Projects A/B/C built end-to-end and sound. 8 API routes at /api/cwp; provenance chain verified all the way to doc_id + PDF page; taxonomy v2 independently re-validated (no SIC overlaps).
  • Gates still open: crosswalk 6/175 rows approved (A.5 unfinished); owner policy decisions of 2026-08-08 scripted in scripts/apply-policy-decisions.js but NOT yet applied to DB; taxonomy v2 drafted but not loaded/frozen/activated; materiality floor still placeholder; all 412 conflicts publishable=false (correct while gates open).
  • Committee membership covers 110th–115th Congress (2007–2018) only — conflicts are absent (not wrong) for 2019+.
  • A concurrent agent session was implementing the crosswalk confidence-floor rule (materiality.json + compute-conflicts.js) in the same tree on 2026-08-08 ~15:45 UTC.

THREAD 1 (pilot selection): NOT started, contrary to prior assumption. No pilot list exists anywhere; the 5 cards in data/conflict-report-cards/ are Project C's technical demo, not a selection. The data-driven selection query works (documented + reproducible in the plan doc, §5): 155-member eligible pool after filtering (≥5 wealth years, zero under-review years, zero obfuscation flags), party-balanced (103 D / 105 R / 2 I). Key discovery: only 202 disclosures corpus-wide are verified_correct and the stark-growth candidates have ~0 verified filings each — so Thread 1 is rank → shortlist (~25) → judgment pass → TARGETED verification of shortlisted members (~1–2 hrs each, anchor-assisted) → lock config/pilot-members.json. Top candidates by absolute growth: Peters D-CA $90M→$213M, DelBene D-WA, Schneider D-IL, Meuser R-PA, Ron Johnson R-WI. By multiple: Zinke R-MT ×32, Newhouse R-WA ×7, Crapo R-ID, Timmons R-SC, Lieu D-CA ×6.

THREAD 2 (report card + pilot pages): had no plan anywhere; now planned in the same doc. Two-level card (crude headline + explanation ladder: raw growth → WGS → ownership pattern → committee overlap → v2 slots), caveats as first-class block, PDF click-throughs everywhere. Hand-assembled template-shaped pages against /api/cwp (owner's lean confirmed; WebBuilder generalization decided AFTER pilot, B7). Deployment: stage on VPN-only dev; public launch hard-gated on materiality floor + crosswalk approval + pilot verification + election-law counsel + corrections workflow + prod security audit + domain cutover (all existing owner tasks).

BUG FOUND: /api/cwp/members/:bioguide/wealth-summary includes under_review snapshot years in baseline/latest/WGS — violates the "under-review years are publicly excluded" principle. Must fix before any card renders from it (plan item T1.0.4/B1).

Show more
GTM breakdown + conflict-pipeline spec → handed to Claude Code
2026-08-06
Edit

Session arc: GTM planning → Report Card model → conflict-detection pipeline spec → execution plan handed to Claude Code.

  1. GTM BREAKDOWN. 18 tasks filed across six workstreams (Identity & Legal, Product Readiness, Website & Landing Pages, Earned Media, Owned/Social, Influencer/Partner, Launch & Measurement). Two decisions resolved: v1 metric = crude before/after (Wealth Growth Score); launch shape = curated pilot of 5–15 stark, bipartisan-balanced cases (not all-435). Full doc: CWP-GTM-Breakdown.md in the projects repo. Cross-project reality: WebBuilder (#1) cannot auto-spawn landing pages yet; SocialPoster (#12) is embryonic — CWP is its first live test. Both need real build work.

  2. TWO-LEVEL REPORT CARD MODEL (owner's framing). Headline = crude "$X to $Y over N years." Underneath = the explanation layer: incoming baseline, trading-behavior onset (few trades year 1 → pattern later), how wealth grew, assets overlapping committee assignments, and market-expected growth. The explanation layer is the evidence that licenses the crude headline; it is also the defamation defense. Inference ladder (weak→strong): raw growth < growth-beyond-market < trading-behavior change < committee/asset overlap with timing. Trade-timing analysis is v2 (needs dated periodic transaction reports + legislative events).

  3. CONFLICT PIPELINE ARCHITECTURE. Core reframe: no fuzzy intersection — committees and assets both normalize onto a shared cwp_sector taxonomy; conflict detection is an exact set join. Fuzziness quarantined to one stage: parsed-asset-string → known-entity resolution (tiered cascade: ticker → exact name → fuzzy ≥0.90 → LLM → manual queue), every row carrying method/confidence/version. Committee→sector crosswalk = ~250 rows incl. subcommittees (the precision layer), LLM-drafted, owner-approved, versioned. Tier 2 broad funds = diversified, excluded. Tier 3 private/LLC = v2. Historical membership sourced from data staged at /srv/app-dev/data/cmtes (~2006-forward; corpus starts 2007, so full coverage). ProPublica Congress API is dead; Congress.gov API is current-Congress only — historical rosters come from the staged files.

  4. API/REPORT LAYER. Acknowledged gap: no reports/endpoints built yet. Spec §6 defines seven v1 read-only endpoints (member committees, assets, conflicts, wealth-summary, report-card, sectors/crosswalk, review-queue). Report Card, WebBuilder pages, and SocialPoster all consume the API — nothing renders from tables. Every payload carries taxonomy/rule/resolver versions + source-PDF pointers (auditability extends into the API).

  5. EXECUTION. Three build projects filed as tasks: A (Taxonomy & Crosswalk), B (Resolution & Conflict Engine, parallel with A until the join), C (API layer & pilot run, gated on A+B). Governing docs in repo: CWP-Conflict-Pipeline-Build-Spec.md + CWP-Conflict-Pipeline-Execution-Plan.md (the latter ends with the Claude Code kickoff prompt; agent's first move is inspecting /data/cmtes and proposing an ingest plan before bulk-loading). Structural safeguards in the build: taxonomy UNFROZEN flag, materiality-floor placeholder with publish-path warning, crosswalk draft/approved flag that blocks publishable conflicts.

  6. OWNER DECISION STATE. #1 taxonomy base = RESOLVED (custom over SIC/NAICS, GICS skipped). Open: #2 crosswalk approval pass (~2–3 hrs, mid-Project-A, gates the join), #3 adjudication surface (recommended spreadsheet in/out; gates pilot run), #4 materiality floor (gates publishing), #5 pilot member list — 10–15, bipartisan-balanced, starkest verified trajectories (gates pilot run).

Show more
Funding Docs
2026-07-19
Edit

Funder Materials — Congressional Wealth Project Two artifacts: the short intro (send cold) and the deeper dive (send on reply). Bracketed [OWNER] items need your numbers/decisions before sending. 2026-07-18.

Quick intro (2–3 sentences, send as-is or trim) Members of Congress must disclose their finances — but the disclosures are designed to be unreadable: vague dollar ranges, scanned handwriting, thousand-page brokerage attachments. The Congressional Wealth Project has turned every House disclosure since 2007 — including the paper-era archive no one else has digitized — into an audited, source-linked record of how much each member's wealth grew in office, and whether that growth can be explained by anything other than the office itself. We deliver the answer where it matters: to the voters of each member's own district.

Deeper dive (send on a warm reply) The problem The STOCK Act made congressional financial disclosure mandatory. It did not make it legible. Values are reported in ranges as wide as "$5,000,001 – $25,000,000." A decade of filings exists only as scans, some handwritten. The wealthiest filers routinely bury their holdings in hundreds of pages of attached brokerage statements. The practical effect: the public record of congressional self-enrichment exists, and almost no one can read it. The services that do parse this data (Quiver Quantitative, Capitol Trades) serve investors — "copy Congress's trades" — not voters, and their coverage begins where digital filing begins.

What we built The complete record, machine-readable. 4,300+ House disclosures (2007–2024) across five generations of forms, parsed into ~170,000 asset records with values, ownership attribution (member / spouse / joint / dependent child), and page-level links back to the source PDFs. The 2007–2012 scanned archive is, to our knowledge, digitized nowhere else. An accuracy system, not just a parser. Every extraction passes automated integrity gates; anything suspicious is flagged for human review rather than published; human-verified reports become permanent test data that every future change to our software must pass. Wealth figures always carry their disclosed ranges, and years still under review are publicly excluded from our calculations. We treat one debunked number as a project-level failure — the entire architecture exists to prevent it. Radical cost efficiency. The engineering approach (deterministic parsing for machine-generated documents, benchmarked AI only for scans) means reprocessing the entire corpus costs tens of dollars, not tens of thousands. Marginal cost per district, per year, approaches zero. What makes it different Wealth-level, not trade-level. Others publish trade tickers. We compute each member's net-worth trajectory across their entire tenure — and, in our next phase, the excess of that growth over what markets explain, ownership patterns (wealth parked with spouses and dependents), an opacity score (filers who make auditing hard are themselves a signal), and overlap between holdings and committee jurisdiction. Archival depth. Pre-2013 filings are where many careers' wealth stories begin. We have them; the trade-feed services don't. District-first distribution. We are not building a national outrage list. The product is a District Report Card — "here is what your representative's disclosures show, with every claim linked to the government document it came from" — delivered through local landing pages, local newsrooms (pre-verified data journalism, free to them), and locally targeted amplification. National lists preach to the converted; district delivery reaches the only voters who can act on it. Auditability as identity. Every published number links to a PDF page. Our methodology, its limits (exempt primary residences, categorical ranges, unknowable pre-office baselines), and our corrections policy are published alongside the findings. Where the project stands Data foundation complete and verified; final pre-publication audit in progress; analytics roadmap (excess growth, committee overlap, live filing alerts) approved and phased; distribution build next. [OWNER: add team/org status — e.g. fiscal sponsorship or 501(c)(3) status, who you are.]

What funding does [OWNER: fill amounts.] Funding accelerates three things: (1) completion of the analytics layer and District Report Cards for all 435 districts; (2) the local distribution engine — landing pages, newsroom partnerships, and legally reviewed local amplification; (3) sustained operations: each new filing season flows through the system at near-zero marginal cost, so funding buys durable public infrastructure, not one-off reports.

Integrity commitments a funder should hold us to Nonpartisan by construction (the same math runs on every member of every party); no publication of unverified data; ranges and limitations always disclosed; corrections policy public; election-law review before any paid targeted communication.

Show more
Internal Operating Doc
2026-07-19
Edit

Internal Operating Brief — Congressional Wealth Project The one-page-per-piece map of what exists, why it exists, and what it's worth. For orientation of any new collaborator (human or agent). 2026-07-18.

Mission Determine, for every member of the U.S. House, whether their wealth grew in ways that office alone cannot explain — and deliver that answer to the voters of their own district. Not a national outrage list; a local accountability instrument, district by district.

Operating principles (these decide arguments) Bad data is worse than missing data. Every number must survive audit; anything uncertain is flagged, ranged, or excluded — never silently shown. Deterministic first. Machine-generated documents are parsed with geometry and code; AI models are reserved for scans and handwriting, and only behind measured benchmarks. Human judgment is ground truth. Corrections and verifications are never overwritten by machines, and every verified report becomes test data that future parser changes must beat. Spend is metered. All model usage runs through cost guardrails with preflight estimates and hard ceilings; vendor spend requires owner approval. (The entire corpus reprocessing in July 2026 cost ~$42.) One source of truth per topic. Routing: PARSING-STRATEGY.md. Accuracy: ACCURACY-REPORT.md. Launch: PRE-LAUNCH-CHECKLIST.md. Roadmap: ENRICHMENT-ANALYTICS-AND-DISTRIBUTION-PLAN.md. The pieces

  1. Ingestion + parsing (the foundation) — WORKING Downloads and parses every House annual financial disclosure, 2007–2024: 4,300+ documents across five document generations, including the scanned and handwritten 2007–2012 archive that exists in structured form nowhere else. Per-form pipelines: Form 4 (modern) via a deterministic coordinate parser (no AI, free, ~5 minutes for the whole corpus); Form 3 (old scans) via benchmarked vision extraction; Forms 1/2 via OCR-table extraction; Senate eFD via HTML. ~170K asset records with values, ownership codes (self/spouse/joint/dependent-child), and page-level provenance. Value: the moat. Everyone else's congressional data starts where the digital era starts; ours reaches back through the paper era, normalized.

  2. Accuracy machinery (the credibility engine) — WORKING Every parse passes invariant gates (impossible values, column confusion, attachment gaps, ordering violations → review queue, never silent trust). Golden corpora of human-verified reports gate every parser or model change. Cross-year anomaly detection catches value misalignments. Wealth years still under review are excluded from headline calculations and publicly footnoted. Value: this is what makes the project publishable. One debunked claim in one district would poison all 435; the machinery exists so that never happens.

  3. Review tooling (the human-leverage layer) — WORKING Admin panel with: anchor-assisted review (a verified year audits its neighbors — divergences, new/missing assets, one-click corrections, saved progress), one-click row edits, page-level PDF jump, verification that instantly recalculates the member's wealth. Attachment-style filings (wealth hidden in appendices) are flagged for human resolution by policy. Value: collapses thousands of row-audits into hundreds; converts owner review time into permanent test data.

  4. Wealth analytics (the product brain) — WGS live; rest planned Live: per-member wealth snapshots (with honest min/max ranges) and the Wealth Growth Score (growth minus reported income). Planned (approved plan, phased for agent execution): excess growth (growth beyond what the market explains — the headline metric), cohort percentiles, an opacity score (filers who make auditing hard are themselves a signal), per-owner breakdowns (assets parked with spouses/children), historical trading extraction from Schedule B, committee-jurisdiction overlap, live PTR tracking for freshness. Value: converts "who is rich" into "whose wealth is unexplained" — the difference between a leaderboard and an accountability instrument.

  5. Distribution (the point of it all) — designed, not yet built District Report Cards (every claim linked to its source PDF page), 435 SEO landing pages (via the separate website-builder project), press kits for local newsrooms, and eventually geo-targeted amplification — hard-gated behind election-law counsel. Triggered by events (new filings), not a constant national drumbeat. Value: "your congressman got rich in office" delivered to his own voters is a different category of thing from a national corruption list.

Current status (one line each) Data: all form types on final pipelines; corpus reparsed; wealth recalced. Launch gate: owner hand-audit + attachment resolutions in progress (PRE-LAUNCH-CHECKLIST.md). Roadmap: ENRICHMENT-ANALYTICS plan approved; executor-agent ready; reviewer gates at every phase.

Show more