Case Studies4 min read·Last verified: June 11, 2026

Case Studies: What the Winners (and the Puzzles) Actually Do

In short: Ten company teardowns from our 18-product audit. Each is built on verbatim evidence: what they shipped, the numbers it produced, what to copy, and what only works because of who they are. Three breakout stories, three engineered mechanisms, two established-player puzzles, two live tensions. Snapshot date 2026-06-11.

The playbook tells you what to do. These pages show you who already did it, with receipts. Every claim traces to our audits and the Data Room. Every page ends with a "what NOT to over-copy" section. The fastest way to waste a quarter is imitating an established player's privilege instead of a challenger's mechanics.

The breakouts: products that won without a training-data head start

Better Auth: #2 docs source at two years old No head start, mediocre trust score, more agent retrieval traffic than React. We break down how they got in, and which part of it you can actually copy.

DodoPayments: taking the "payments" query Stripe never claimed Description engineering in its purest form. The registry description states the task, so the task query returns them, not the established player.

Resend: the smallest corpus with the highest score Benchmark 92.3 on a corpus a fraction of its rivals' size, plus the only .well-known agent-discovery manifests we found in the wild.

The mechanisms: engineered loops you can copy

Bun: the growth hack that is a filename bun init writes use-bun-instead-of-node-vite-npm-pnpm.mdc into every project. It is disclosed, opt-out-able, and measured. The mechanism flipped agent choice 100% in our pilots.

Next.js: what running every play at once looks like Docs inside the npm package, AGENTS.md by default at scaffold time, content negotiation, a public benchmark. The reference build that does it all.

Convex: the eval-driven operator Rules tuned "using rigorous evals," a public LLM leaderboard, and a CLI that keeps customers' AGENTS.md current. Measurement as the strategy, not the afterthought.

The established players: privilege, and what it costs

Stripe: fighting its own training-data ghost The 265K-snippet docs index and an llms.txt that orders agents to never recommend Stripe's own legacy API. The established player's problem isn't absence. It is being confidently wrong in the model's memory.

The Tailwind Puzzle: top-10 with zero agent surface No llms.txt, no markdown, no MCP. They are carried by training-data mass and third-party volunteers. We show why that path is closed to you, and where it leaks even for them (agents emit obsolete v3 config 100% of the time).

The tensions: live experiments worth watching

Clerk: docs written for agents, not developers The strongest instruction layer we audited, paired with the weakest retrieval-quality score among winners. Instruction vs retrieval investment: the natural experiment is running now.

Drizzle: 440 snippets that beat 2,297 (with Polar and Hono as contrasts) Minimal-but-dense beats full-stack-but-stale (Polar) and strong-but-uncurated (Hono). Density is the rubric.

Which case study for which play?

If you're working on…Read
Registries & description engineering (Play 2)DodoPayments · Drizzle
llms.txt & directives (Plays 5, 8)Stripe · Clerk · Tailwind (the cost of skipping it)
Markdown & snippets (Plays 6, 7)Next.js · Resend · Drizzle
MCP & skills (Plays 3, 4)Better Auth · Resend · Convex
Scaffolders & environment (Play 9)Bun · Next.js · Convex
Evals & leaderboards (Play 11)Convex · Next.js

How to read these

All ten are point-in-time teardowns (2026-06-11; single-day metrics carry ±10% error bars) of surviving winners plus two deliberate counterexamples. Survivorship caveats are stated in each file, and items we couldn't verify are flagged UNVERIFIED rather than asserted. They are anecdotal evidence by design. The systematic data lives in Part 2 and the Data Room. These pages supply the texture the aggregates can't.

FAQ

Are these endorsements or paid placements? Neither. Companies appear because they showed up in our research data; none were contacted, none paid, and the counterexamples didn't volunteer. Critique and praise both trace to published evidence.

Will these be updated? Quarterly, with the rest of the guide. Case studies are re-snapshotted and changes logged in the Data Room. A teardown that ages badly is itself a finding (see openclaw's −50% month in Part 2).

Why isn't [company] here? The set covers the 18 products in our original audit, selected for lesson density. Suggest additions via the contact page. Products with a measurable, distinctive mechanism get priority.


Last verified 2026-06-11. We re-test the claims on this page quarterly. Changes are logged in the Data Room.

Part of The Complete Playbook to Agentic Discovery.

Stay ahead of the agents. We re-test this playbook quarterly and publish what changed: new data, busted myths, ranking shifts. Get the update digest →

Want this done for you? Synscribe runs agentic-discovery programs for B2B SaaS and developer platforms. Talk to us →

Get a diagnosis

Are you the default an AI agent reaches for?

Get an agent-readiness diagnosis of your product, especially API and MCP products, plus the punch-list to become the one agents pick by default.

Subscribe to research