Fractl ResearchAugust 2026~5,000 ideas · 20+ methods · 21,000 judgments · 1,200+ blind human labels from 8 readers

Escaping the idea basin

We asked an AI for 40 campaign ideas and got 14 — the rest were the same ideas re-worded. So we spent $110 testing every published fix, up to and including surgery on a 72-billion-parameter model's weights. What actually widened the ideas wasn't what anyone predicted.

The bottom line, up front

1. Ask a frontier model for 40 ideas and you get about 14. One premise appeared seven times in a single batch. The collapse is baked into the model by preference training, not a settings problem.

2. The famous knobs don't work. Temperature: nothing. "Verbalized sampling": nothing. The tricks that do widen the set — spec decomposition, personas, provocation prompts — pay for most of that width with ideas no strategist would run. A seven-strategist panel later found one partial exception: spec decomposition's surviving spread is real on our main brief (details in the viability section).

3. AI judges cannot score originality. Four frontier models agreed with a blind human read 43–55% of the time. Coin flips.

4. Rewriting the brief — naming the goal, licensing tangents, requiring a through-line — nearly doubled the usable, distinct ideas from a plain prompt: 12 → 20. The single biggest lever we found.

5. Different model lineages have different favorite ideas. Pooling four generators tripled the distinct viable ideas (21 → 64) with almost no overlap between sources. The widest swath comes from a portfolio, not a better prompt.

Status: v1.1 — panel read in; replication of findings 4–5 in progress

Update, September 1: seven more Fractl strategists have now blind-read the plain vs. spec-first batches on both briefs. The viable-rate gap replicated; one yield conclusion moved and is logged in the corrections log below. Findings 4 and 5 still rest on one senior strategist's blind reads and one replicate; that round is still running, and we'll update this page either way.

The basin is real, and deeper than you think

Every aligned language model has favorite ideas. Ask for 40 earned-media campaign concepts and 18% of them are the same idea — "humans vs. AI, blind test" — wearing different headlines. 57% sit in one emotional register. Clustered strictly for "same idea," a 40-idea batch holds roughly 14 actual ideas.

This isn't randomness failing you. Preference training teaches models that typical answers are good answers — researchers call it typicality bias — so the model learns to stay near the middle of everything it knows. The 2026 literature is close to unanimous that the collapse lives in the weights, and our measurements agree:

Wider isn't better: the viability trap

We had a strategist blind-read whole 40-idea batches and tick every idea he'd actually take to a client. Plain prompt: 33 of 40 viable. The spec-first "winner": 22 of 40. Cluster the survivors and both land near 12 usable, distinct ideas. The diversity machinery bought its spread almost entirely with ideas a human throws out.

Almost every method that widened the set did it by escaping the brief instead of escaping the basin.

A "provocation" instruction — render each idea as the strangest version of itself — made the trap vivid: machine-scored originality soared while human-scored viability collapsed to a 0.28 win rate. Surprising and un-runnable is not a campaign.

Update (September 2026): seven more readers, one conclusion moves

Because this result carried so much weight, we re-ran it as a panel: seven Fractl strategists blind-read the same plain and spec-first batches on both briefs — same sheets, same tick-every-viable instruction, no knowledge of which set was which. The viable-rate gap replicated cleanly: plain wins on usable ideas everywhere (readers averaged 23 vs 20 of 40 on the main brief, 29 vs 22 on the second). The yield conclusion moved. Keep only the ideas a majority of the panel ticked, cluster the survivors:

Majority-vote survivors (7 readers)ViableDistinct & viable
Main brief — plain prompt25/4012
Main brief — spec-first18/4015
Second brief — plain prompt27/4010
Second brief — spec-first18/409

Majority = ticked by ≥4 of 7 readers (main brief) / ≥4 of 6 (second brief). Distinctness = strict same-idea clustering of the surviving ideas.

On the main brief, spec-first's survivors span 15 distinct ideas against plain's 12 — a modest real edge, not the wash a single reader saw, and it holds for six of the seven readers individually; on the second brief it's a wash. The mechanism is arithmetic, not taste: plain's 40 ideas only contain 14 distinct ideas to begin with, so its 25 survivors pile into the same clusters, while spec-first's 20 starting clusters survive a harsher cut with more of them still standing. Spec-first is still the wasteful way to widen a set — a quarter of its ideas die at review — but the spread that survives review is real spread. "Flat or worse" becomes "flat to modestly better."

The judges have their own basin

We tried to automate the human read. Four frontier models from three companies judged ~21,000 blind head-to-head pairs. Checked against 139 decisive blind human picks across two briefs:

Judge signalAgreement with human
Pairwise originality (Bradley-Terry, 4 judges)43–55%
Pairwise pitchability41–51%
Embedding nearest-neighbor novelty58%
Rubric quality score39% — anti-correlated

Agreement measured on decisive (non-tie) pairs only. 50% = coin flip.

The failure has a mechanism. Judges score pairs in isolation and never saturate: a human reading 60 pairs is tired of "the AI writes its own obituary" by pair 12; the judges rewarded it every single time. In one experiment, 13% of all judge-favored ideas on a moving-company brief were the same divorce-data premise. And the judges share the generators' training, so they share the generators' taste — a judge can't see its own blind spot.

One technique helped: showing the judge the candidate pool first and declaring its premises used up lifted originality agreement from ~50% to 63% — and to 77% on the pairs where two judge families agreed. Useful for triage; not good enough to replace a human. Every ranking in this study that matters was made by a person.

The $110 brain surgery

If collapse lives in the weights, edit the weights. Open-weight models ship in pairs — the raw pretrained "base" model (the whole distribution, unaligned) and the aligned "instruct" model — with identical architecture. So you can blend them, parameter by parameter. We interpolated Qwen2.5-72B at 30 / 50 / 70% base on a rented H200 and measured how many regions of idea-space each blend reached:

ConfigurationRegion reach (Hill number)
Instruct (0% base)8.1
30% base10–11
50% base9–12
70% base13–17
Plain Claude Sonnet, no tricks26.4

The dial is real. Every step toward base widened the ideas, monotonically, and the model still followed instructions at 70% base. It's a lever no API will ever expose. And it wasn't enough. The open model's basin was so much deeper than the frontier model's that maximum blend never caught a plain Sonnet prompt on the main brief.

Two more findings from the open-weights lane:

The lever nobody was testing: the brief

Halfway through, our strategist said the wide-model ideas were "cool, but I'm not sure what they're trying to be about." Our brief said ideas "must be something the company could produce and pitch to journalists" — and every generator read that differently. So we rewrote it the way he actually judges: name the goal, license tangents explicitly, but require a through-line — a CTA or mention of the product must not feel out of place in the finished project.

Plain prompt, same modelViableDistinct & viable
Original brief33/40~12
Rewritten brief38/40~20

No new model. No technique. A clearer assignment nearly doubled the usable yield — more than anything else we tested, machinery included. A vague brief collapses the model into its safest cluster; a brief that names the goal and licenses tangents gives it permission to spread and a tether that keeps the spread useful.

No single generator wins. The portfolio does.

The final surprise. Under the rewritten brief, four very different generators each produced roughly 12–21 distinct viable ideas — call it a tie. Then we pooled everything the human had ticked viable and clustered the pool:

SourceAdds to the pool
Plain frontier prompt21 distinct viable ideas
+ Spec-first decomposition+15 new
+ 70%-base open-weights blend, anchored+14 new
+ De-aligned 70B harvest, anchored+15 new
Union64 distinct viable ideas — only 7 of 64 classes overlap

Different model lineages mine almost completely different veins of good ideas — same brief, same bar, near-zero overlap. Three times the swath of the best single method. The machinery earns its place not by beating the frontier model but by being uncorrelated with it.

The playbook

First step, today: take your next ideation brief and add three sentences — the goal, permission to be tangential, and the through-line rule. Then ask two unrelated models, not one model twice.

Method & limits

Scale: two fixed briefs (our own AI platform; a national moving company) · 15 prompting conditions × 3 replicates × 40 ideas × 2 briefs in the main round, plus open-weight and de-aligned-model conditions (~5,000 generated ideas total) · ~21,000 head-to-head judge calls · 230+ blind human-labeled pairs and six whole-set human reads. The measurement grid (subject × mechanism × register) was declared before any generation. Spread measured by strict same-idea partitions under two judge families, grid Hill numbers, and embedding Vendi scores; where the judge families disagree we report the range. Open weights: Qwen2.5-72B base + instruct, linear weight interpolation, fp8 serving on a single H200; total GPU spend $9.50. Limits: two briefs in one broad vertical; viability is one senior strategist's blind judgment (additional readers in progress — the two flagged findings are gated on them), not market outcome data; the blend's premises were harvested under an earlier brief wording than the hosted models' (the anchoring render mitigates this; noted). Panel read (added Sept 1): the plain-vs-spec-first viability result now rests on seven strategists' blind reads across both briefs (~1,040 idea-level viability labels; inter-reader agreement kappa .28–.56; majority vote for the reported numbers). Findings 4 and 5 remain single-reader until their round completes.

Corrections log: (1) An early condition hard-coded one brief's subject into the other brief's prompt; regenerated and rescored. (2) Our first human-read sheet sampled judge-favored ideas — circular; resampled uniformly, re-read. (3) An early "86% judge agreement" figure came from 14 pairs and did not survive 139; retracted. (4) One model server returned identical "replicates" from a fixed seed; regenerated with per-request seeds. (5) Sept 1: v1 said spec-first's distinct-viable yield was "flat or worse" than a plain prompt, based on one reader. A seven-strategist panel replicated the viable-rate gap but moved the yield direction on the main brief — majority-vote survivors span 15 distinct ideas for spec-first vs 12 for plain (wash on the second brief). Softened to "flat to modestly better"; the waste finding stands. All model calls are cached and reproducible.

Disclosure: Fractl is a marketing agency; the ideation layer this research informs runs inside Fractl Agents, which we build and sell. Weigh the piece knowing the interest exists; the method is open and the numbers are checkable.