Progressive Labs
CONTEXT Current tool set: Ask (open-ended Q&A on corpus), Compare (top-N tactics for your circumstance), Campaign Planner (coming soon — full plan). User feels a tool is missing and proposed a "Validation" tool (aka bullshit meter / article analyzer): user pastes a pre-canned efficacy statement, a blog/article URL, or a raw claim (e.g. "SMS has a pp of 3.5pp"), and the tool describes where it's right, wrong, and the research around it. Core problem it targets: research skews toward low-salience elections, which inherently produce higher pp effects; those small-campaign/low-salience results get over-extrapolated into generally-accepted "knowledge."
ASSESSMENT OF THE VALIDATION TOOL
- Strategically the best-sequenced third tool — better than Campaign Planner. It's descriptive/epistemic, zero prescription, low blast radius. Fits the standing discipline of "don't drift into Plan" (Plan issues prescriptions we can't yet back with our own validated outcomes).
- It's also the most direct expression of the moat: the credential thesis is that Ask must know the LIMITS of the literature. This tool turns that credential into its own front door — the single feature that most separates PL from a GPT wrapper.
- Distinct job from Ask (not just a mode): Ask = "I have a question, what's the answer." This = reverse vector, "I found a claim in the wild, adjudicate it for me." URL ingestion meets real behavior (people forward each other decks/blog posts).
TWO SHARPENINGS
- Kill the "bullshit meter" framing. Users ARE the practitioners we're auditioning for track-two relationships, and many are repeating these over-extrapolated numbers in their own fundraising decks. A tool that feels built to call them wrong reads as adversarial. Goal = they walk away feeling MORE competent, not corrected. Same mechanism as Ask.
- Reframe from truth to transportability. The problem isn't a truth problem, it's a transportability problem. "SMS = 3.5pp" is usually TRUE in its own study; it breaks only when transported to a different context (higher salience, higher baseline turnout). So the tool's job = take (claim) + (user context: race type, salience, baseline turnout) and return: the effect it rests on, the context it came from, how much it likely shrinks/holds for you, and the uncertainty. Rename to Applicability / "Does-This-Travel." (Note: Applicability = the user's proposed Validation tool, renamed. NOT the evidence-gap map, which is separate.)
OTHER IDEAS WEIGHED AGAINST IT
- Test/holdout designer (highest ceiling): the "make it trivial to design a measurable program — universe split, randomized holdout, pre-registered analysis plan" idea. This basically IS the Analyst Institute partnership / the experiment-design + meta-analysis layer, and the only path to proprietary data vs. re-serving the public corpus. More valuable long-term but relationship-gated — not a v1. The free tools earn the right to build it.
- Evidence-gap / coverage map (sleeper, separate tool): systemic version of the same core problem. Instead of debunking one claim at a time, it visualizes where the evidence base is thin (which tactics, race types, salience levels have real RCT support vs. one small low-salience study extrapolated everywhere). Cheap off existing corpus, doubles as the funder pitch ("evidence base for local races is thin — here's proof, fund N tests"), and is a standing credibility asset rather than a one-off utility.
RECOMMENDED SEQUENCE
- Build the Applicability tool (the user's Validation tool, reframed to transportability) as the third shipped tool — low-risk, and it IS the credential.
- Prototype the evidence-gap map alongside it — shares the backend, carries the funder story.
- Hold the test/holdout designer until the relationships are real.
Design the first tests to need almost nothing from the campaign. "Execute a test with us" sounds like asking a stressed campaign manager to co-run a science project. But the highest-value first experiments are holdout slices of programs campaigns are already buying: they're sending 20,000 social-pressure mailers anyway — the ask is "let us randomly withhold 2,000 and we'll tell you, from the public voter file after the election, exactly what your mail bought you." That's not a favor to you; it's free program evaluation for them. Zero budget risk, zero execution change, and turnout outcomes are administrative data you can collect without their cooperation after election day. Trust required: minimal. This works before you have any authority.
Borrow credibility instead of building it from scratch. Two parties are hungry for exactly what you have. Academics need field sites — the entire GOTV literature exists because researchers found campaigns willing to lend their programs; a PI partnership gets you IRB cover, free analysis labor, and instant methodological authority in front of funders. Vendors (relational apps, texting platforms) desperately need independent validation to sell — they'll co-fund tests and recruit their own campaign customers for you. Notice both of the gray-lit decks we found in your corpus were vendor/funder-commissioned tests kept confidential — publishing yours openly is itself the differentiator.
Bring funders a portfolio, not a demo. When you do go to funders, the pitch isn't "our tool is good" — it's "the evidence base for local races is thin (we can show exactly where — our own coverage data proves it), and for $X we'll run N randomized tests across M city-council and state-leg races in one cycle, published openly." Small races are where tests are cheapest, where the literature gaps are largest, and where your tool already scored 7/10. Funders fund designs, and you're uniquely positioned to show up with one — including the audit trail from this review as proof you grade your own homework.
Session · 2026-06-12
That read is right, and the sequencing does more than just stage the rollout — it solves your "not enough to attract AI yet" problem directly. Ask and Compare are the things you can build and ship alone, with no partner, no data-sharing agreement, no org buy-in. They run on the public corpus. And a working tool that practitioners actually use is a far better thing to walk into AI with than a pitch deck. You'd be showing up having already demonstrated the translation layer works and that people want it — which is the only version of that conversation where you're a peer offering to extend their reach, rather than a supplicant asking for their gated data. Build the asset that earns the relationship; don't ask for the relationship to build the asset.
One sharpening, since you framed these as two tracks. They're not really parallel — track two (relationships, test design, generating data) is gated by track one, and specifically gated by Ask getting good. Here's the mechanism: the orgs you want to design experiments with will judge your competence by how good your tools are. If a program director plays with Ask and it gives sharp, well-caveated, genuinely useful answers about their actual tactics, that's what earns the meeting where you propose running a holdout test with them. Ask isn't just a wedge product — it's your credential. So the build order has a real dependency: Ask has to be good enough to be a credential before track two is even reachable. That raises the bar on Ask. It can't be a thin GPT-wrapper Q&A; it has to demonstrably know the literature and, critically, know the limits of the literature, because the people evaluating it are the few who can tell the difference.
The thing I'd watch as you start: resist letting Compare quietly drift into Plan. The pull will be strong — a campaign will ask "okay, so what should I do," and the natural next feature is to just tell them. But Plan is the high-blast-radius product we flagged, and shipping it before you have the feedback loop means you're issuing prescriptions you can't yet back with your own validated outcomes — only with transported effects from other contexts. Keep the early product descriptive ("here's what the evidence says about these options, and here's the uncertainty") rather than prescriptive ("here's your plan"). That discipline is also exactly what keeps you positioned as AI's complement instead of drifting toward the Movement-Labs-style "we'll build and sell you the tool" posture you defined yourself against.
So if I were drawing the line for the next phase: get Ask good enough that a skeptical research-literate program director comes away impressed, ship Compare alongside it as the natural companion, and treat every one of those early users as a relationship being auditioned for track two — not a transaction. The data and the partnership are both downstream of that, and both are easier to reach than they look from here.
What's the current state of Ask — is it at the "impresses a skeptical expert" bar yet, or is that the immediate work?
Session · 2026-06-12
Settling partner-not-competitor and influencer = the infrastructure layer (advocacy orgs, big donors, party committees) actually tightens the whole strategy, and it does something you didn't quite say out loud: that layer is the one tier of actors with the standing to mandate measurement. A donor can say "we'll fund this turnout program if it runs with a randomized holdout and reports back." A state table can make it a condition of the coordinated plan. Campaigns will never volunteer rigor; the infrastructure layer can impose it. So your influencer thesis and your feedback-loop thesis aren't two strategies — the influencers are precisely the lever that makes the loop possible.
Now the pushback, because I think you've talked yourself into a false binary. You laid out two options: get campaigns to self-report (unreliable), or wire into every execution tool via API/MCP (monumental). Your instinct that the all-tools integration is "probably useless" is correct — but not for the reason you gave. It's not useless because it's big. It's useless because integration gives you correlation, not causation. Even if you had a perfect live feed from every texting platform, dialer, and canvassing app, you'd get "we contacted 50k people and turnout was 61%," which tells you nothing about what the contact did. Without a control group you can't separate the tactic's effect from the fact that you targeted likely voters in the first place. So the firehose of execution telemetry is exactly the wrong thing to chase.
Here's the third path you skipped, and it's the one the field actually uses. You don't instrument the delivery. You control two things only: the assignment and the outcome.
- Before a program runs, define the treated universe and a randomized holdout. That's just two lists.
- The org runs the program in whatever tool they already use — they just exclude the holdout. Zero integration required on the execution side.
- After the election, match both universes against the voter file for the only outcomes that matter and are already centralized: did they vote, did they register. Turnout is public record; the progressive data co-ops (Catalist, The Movement Cooperative — note TMC is on AI's partner list — the VAN/voter-file layer) already aggregate it.
- You compute the lift, tag it with full context, and fold it into the evidence base.
That's a complete feedback loop with no tool integrations at all. The only place an API ever earns its keep is the outcome side — pulling the validated voter-file match — and even that can start as a manual data handoff and get automated later. The "monumental integration across all tools" was a red herring; you were trying to instrument the part of the pipeline that doesn't actually carry the causal information.
Two things fall out of this that should make it feel less impossible. First, this is the AI partnership. What you'd be building — making it trivial to design a measurable program (the universe split, the holdout, the analysis plan) and then learn from it — is the executable layer on top of AI's "experiment-ready" funnel. You stop being the telemetry layer and become the experiment-design-plus-meta-analysis layer, which is the thing AI would actually want a partner to do. Second, you don't need N=all. Meta-analysis needs N=clean, not N=everything. A handful of well-designed, well-matched programs per cycle, contributed by your influencer orgs, compounds. The asset is a slowly growing, well-curated body of "here's what we learned this cycle, in context" — not a real-time dashboard of the whole movement.
The honest caveat: the binding constraint here was never technical, it's social. Orgs are protective of their data and their "what works" story is part of their fundraising narrative — getting them to contribute into a shared substrate is a turf problem, which is the whole reason TMC was hard to stand up. But that's exactly why you build the relationships first, with the free translation tools (Ask/Compare), and earn the loop later. The loop is a three-to-five-year compounding asset built on trust, not a v1 feature you need to crack now.
One reframe worth naming: "serve the donor/party layer, no monetization yet" basically describes a funded movement-infrastructure nonprofit — which is AI's own model (donor-funded, membership free). If that's the shape, your moat isn't a product feature at all; it's becoming the blessed, designated translation-and-measurement layer of the ecosystem. That changes what you optimize for.
So the question I'd put back to you: are you willing to make experimental design — the holdout, the pre-registered analysis plan — a first-class part of the product rather than an afterthought? Because if the answer is yes, the loop stops being nearly impossible and becomes mostly a relationship-sequencing problem. If it's no, you're back to self-report and the loop really is dead.