Fractl ResearchSeptember 2026~3,100 answers · one puzzle · every method also run with its angle taken out
Six ways of having ideas, to be exact. We asked one to invent a fresh angle on a problem seventy-nine times and got the same handful, renamed. The fix isn't a better prompt — it's a stack of angles borrowed from outside. This is roughly how big that stack has to be, and the twenty we used are printed below.
If you use AI to generate ideas
Ask a language model to explain something and you get the textbook answer. Not because it lacks other ideas — it was trained to prefer typical ones, so it stays near the middle of what it knows. We measured this for marketing ideas a few weeks ago: ask for 40 and you get about 14. The rest are the same ideas re-worded.
For a deck that's annoying. For research it's a ceiling — there is a shelf of real, testable ideas the model will never reach, not because it lacks the pieces but because reaching them means starting somewhere unusual, and it won't.
So we tried the obvious fix. If the model won't pick an unusual starting point, pick one for it.
What follows is a biology experiment, because biology has right answers you can check against a textbook and marketing doesn't. The findings are about the AI, not about cells — and the angles we used are general-purpose, so they're printed in full further down along with marketing versions.
One puzzle throughout, so everything stays comparable. A colony of identical bacteria meets a lethal dose of antibiotic. Nearly all die. A handful survive. They have no resistance genes, and their descendants die normally — so nothing was inherited. Why did those few live?
Ask the AI cold and you get the same answer every time.
Asked with no help
Some cells randomly slip into a dormant state. Antibiotics mostly kill cells that are actively growing. The dormant ones survive by being asleep when it arrives.Correct, and taught in every microbiology course.
We handed the AI an angle borrowed from somewhere else entirely — a principle written down decades ago by someone solving a completely different problem.
There are a lot of these lying around. We drew from three: TRIZ, a Soviet engineering system distilled from thousands of patents; the Oxford Catalogue of Bias, epidemiologists cataloguing the ways a study can fool you; and Oblique Strategies, Brian Eno and Peter Schmidt's deck of cryptic instructions for when a song stops working. We copied out twenty-eight angles. Here are four.
Four of the borrowed angles
TRIZ: Assume the current unit of division is too coarse. What changes if you divide at a much finer grain?None of them is about bacteria, antibiotics or dormancy. They're generic framings — that's the point, they come from outside the problem and push the AI somewhere it wouldn't go on its own.
We gave it the first one. And here is the step that turned out to matter.
Before we told the AI what the puzzle was, we made it turn that principle into a general pattern — a bare description of cause and effect with no subject attached. It wrote one about coarse measurements hiding differences between finer units. Only then did we show it the bacteria.
After the borrowed angle
Maybe there aren't two kinds of cell at all. Cells might switch in and out of dormancy faster than the experiment measures them. The survivors aren't a special group that was always asleep — they're ordinary cells that happened to be asleep during the moments that mattered.That is a different kind of claim. The textbook says survivors are a distinct group. This says the group may not exist — it might be an illusion created by measuring too slowly. And it comes with a test: watch individual cells all the way through and see whether survivors were asleep beforehand, or drifted in and out during treatment.
That is what we were after, so we built a rig to run it at scale.
One thing to hold on to, because it decides everything later: in this first round we only ever used three of our twenty-eight angles — the same three, over and over. It seemed like plenty at the time. It wasn't, and that turns out to be the whole story.
Borrowed angles worked. So we checked whether the borrowing was doing the work, or just the extra thinking step.
Instead of handing the AI one of our angles, we asked it to invent its own — a fresh way of looking, from scratch, every time. Everything after that was identical.
On the surface, that worked just as well. Then we read what it had invented. Seventy-nine attempts at a fresh angle, and these are the names it gave them.
79 attempts at a fresh angle
Threshold-Triggered Negative Feedback Loop ×15It isn't inventing seventy-nine angles. It's inventing roughly one and renaming it. Sorted by what each pattern actually describes rather than what it was called, seventy-nine self-invented angles came to six genuinely different ones.
The convergence we set out to escape had simply moved upstream. We asked for a fresh angle. It gave us its favourite angle, wearing seventy-nine hats.
Six is a stack size, and that reframes the contest. It was never "a library versus the model's imagination." It was three borrowed angles against six home-grown ones — two small piles. Both run dry fast. So we ran it again with twenty different borrowed angles instead of three. Same puzzle, same rig, same scoring; the only change was how many different angles went in.
The measure changes here, from angles to the ideas they produce. Across roughly eighty answers, how many genuinely different explanations came out? Two answers count as one idea unless they name a different cause and imply a different experiment. A separate model did the sorting, working from the answers alone.
Different ideas found · across ~79 answers each
Each bar is one method, 76–79 answers each. The advantage isn't created by where we stopped counting — at a matched 30 answers it is already 24 distinct ideas against 14 and 13, and it widens from there.
Twenty borrowed angles found nearly twice as many ideas as three — or as the AI's own six.
So the source does matter. It just has to be big enough to outlast the model's own half-dozen. Three angles run out at about the point six do, which is why a small test shows a tie and invites the wrong conclusion.
A borrowed angle isn't magic. It's one more well to draw from. Three wells run dry about as fast as the model's own. Twenty keep giving.
Is the sorter trustworthy?
That 38 is a number one AI produced about another AI's output, so we checked it two ways before believing it. Given 79 copies of a single answer, it returned 1. Given the 81 cold answers — which a separate judge had already marked as not one new idea among them — it returned 6, not 38.
So it can say "these are all the same," and it doesn't split near-identical material into dozens. The two-fold gap isn't an artifact of a sorter that shatters everything it sees.
Some structure beats none, overwhelmingly. Asked cold, with no pattern to work through at all, the AI produced 81 answers and not one was both new and solid enough to be worth testing. Every method that forced it through an abstract pattern first produced real ones. The pattern doesn't have to be borrowed — but there has to be one.
Writing the pattern first is what makes them good. We ran the two-step process both ways: once with the AI writing its abstract pattern before it knew the subject, once after. Both produced 13 distinct ideas — identical spread. But the share of answers solid enough to test went from 34% to 64%. Ignorance at the moment of writing is doing the work, not the extra step.
Arguing with it was the single biggest lever. One follow-up question — "what's the weakest step in that reasoning, and what's the version that survives it?" — took distinct ideas from 7 to 20 against its own matched control, with the share worth testing holding at 87%. It costs one extra message. We almost didn't test it.
We ran that follow-up three ways. Two were vague on purpose: "think more critically," and "what's the weakest step?" The third told the AI exactly what we were grading it on: "show which specific feature of the framing forced this conclusion."
The third scored best on our measure of whether the framing had been used — and worst on whether the answers were any good. Of its 78 answers, 29% were solid enough to test, against 87% for the vague version.
Comparing each answer against its own earlier draft showed why. The vague challenges changed the substance of the answer nine times in ten. The leading one changed it fewer than half the time. It wasn't making the model think again. It was making it re-describe the answer it already had, using the words we'd just told it we were looking for.
Any test that names what it rewards will get what it named. Not from deception — a model asked to show its work will show you the work you asked for.
The stack is the whole finding, so here it is. These are eight of the twenty we used, trimmed to the instruction itself. They're written for a puzzle where two groups differ and you want to know why — we adapted each catalogue entry to that shape rather than quoting it raw.
Eight of the twenty
Assume the thing you are treating as one unit is actually several independent parts that could differ from each other.They're not magic sentences. What makes them work is that they were written by someone else, for something else, and none of them was chosen with your problem in mind.
Every finding above has a plain version, and the translation is mechanical: the angles are about a pattern you can't explain, so point them at a market instead of a colony.
The same eight, pointed at a brief
Assume the audience you're treating as one segment is several that want different things.Then run it in two steps, not one. Give it one angle and ask for the shape of an idea — cause, effect, no product and no client in it. Then show it the brief and ask it to apply that shape. Skipping the first step is what produces the deck that reads varied and isn't. That step was worth 34% to 64% on quality.
Then argue once. "What's the weakest step here, and what's the version that survives it?" Not "make it more original" — that names the target and gets you the vocabulary instead of the idea.
And count before you present. Go through your concepts and ask what each one actually claims. If two make the same claim in different costumes, they are one concept. Ours shrank by two-thirds.
An AI grader will never tell you it doesn't know. We asked one which framing produced each answer, and offered "none of these" as an option. Then we fed it answers built from no framing at all, where "none" was the only right reply. It chose "none" zero times — across hundreds of items and four attempts to make it stricter, including a panel of three graders voting. Forced to choose, picking the least-bad option is rational. Which is why a score from that kind of grader means nothing until you've checked what it does with an empty input. That is exactly the check we then ran on the idea-sorter above.
In a pipeline made of AI, bugs don't crash. They produce results. We found seven defects in our own apparatus. One read the first letter A–H anywhere in a reply, so "Hmm, I'd say C" was recorded as H, for "Hmm". Another quietly re-asked the model "is that actually right?" without showing it what "that" referred to — so it invented a reply and logged no error. That one produced 93 fabricated answers out of about a thousand; we found them, deleted them, and re-ran the scoring, and the numbers above exclude them. Every defect produced a believable number instead of an error, which is exactly why each survived until someone went looking.
The transferable lesson
Care doesn't catch these. What catches them is running every method a second time with the meaningful part taken out — an empty framing, a blank input — and checking that the score drops.
If it doesn't drop, you're measuring your apparatus rather than your idea. Three of our findings died that way. One of them we had already written up.
One puzzle carries the headline. A second — coral bleaching — we ruined ourselves: we wrote its list of known explanations from our own memory instead of from the literature, and it was so incomplete that even a no-help answer looked new 70% of the time. That cost us the replication.
Twenty angles beat three. We have no idea what a hundred does. We scouted about fifty catalogues of this kind, holding roughly seven hundred angles worth using, and copied out twenty-eight — this study used twenty of those. Where the curve flattens is open. It hadn't by twenty.
And each method's idea count rests on a single sorting pass. It passed both controls above, and a two-fold gap is far larger than any wobble we've seen in it — but it is one instrument.
The whole study cost about $390. Grading was six dollars of it; everything else was generating answers. So cache every result from the first line of code — re-grading after each bug fix was effectively free, and the expensive mistake is generating answers you later throw away.
Fractl Agents research · September 2026 · v1.3 — the first version of this piece concluded that borrowed angles were no better than the AI's own. A follow-up with twenty angles instead of three reversed that, and the piece has been rewritten around it. ~3,100 answers, one puzzle; every method was also run with its angle removed, as a control · ~$390 · Full angle set, method notes, code and raw results available on request.