Design · 10 August 2026News

Claude Fable 5 + a Free Design Skill = $10,000 Websites. Here's What Actually Does the Work.

Open the skill and three real mechanisms come out: it interviews you before building, it uses the expensive model only where it shows, and it argues with the plan before there is anything to argue about. All three work without installing anything.

Asfandyar Malik
Long-form resourceEvidence-backed researchWatch the reference video

A free skill plus Fable 5 makes $10,000 websites. Here is what is really doing the work.

The pitch is a free design skill, dropped into Claude Code, that turns one prompt into the kind of site an agency charges five figures for. The line that sells it is that the driver and the vehicle both matter — it is not just the model, it is the skill you drive it with.

That line is true, and it is stated in a way that makes the two sound equal. They are not, and it is worth knowing which is which before you decide the skill is the magic part.

First, the $10,000. That is what agencies charge for a site. It is not what the output is worth, and nothing about a generated page changes an agency's rate. Treat it as a description of the market being disrupted, not a valuation of what comes out.

Second, and more usefully: these skills are text files you can open. So instead of guessing, you can read what is in one. When you do, three real mechanisms come out — and none of them is taste.

Anthropic's Claude Fable product page.
The vehicle. The driver is a markdown file, which is why you can check what it actually does.

Checked 10 August 2026. Model dates and prices come from Anthropic's own newsroom and pricing page, linked at the end. This page describes a test and does not publish results from a run that has not happened — see the limits section.

Three mechanisms, and you can read all of them.

A skill is a markdown file with a settings block at the top, sitting in a folder the agent reads. No code runs. So "the skill" is a set of instructions — and in the good ones, the instructions do three specific jobs.

  1. It interviews you before it builds. Rather than acting on your one-line request, it asks questions first.
  2. It routes work between models. The expensive model does not have to do every step.
  3. It argues with itself before it starts. A deliberate critical pass at the beginning, not at the end.

Those are three genuinely useful ideas and they have nothing to do with the skill having good taste. They are process. You could apply all three by hand. The skill's value is that you do not have to remember to.

Worth taking each one seriously, because they are the transferable part.

It asks you questions before it writes anything.

You drop the file in and say you would like to run the skill, and the first thing that happens is not a website. It is a setup conversation — what are you building, what mode do you want, where should the images come from.

This is the single highest-value thing in the whole category, and it is almost never the part people highlight, because a video of an agent asking questions is boring.

The reason it works: the usual failure of one-shot generation is not that the model lacks skill. It is that "build me a landing page for my consultancy" contains almost no information, so the model fills the gaps with the average of every landing page it has seen. That is exactly what generic AI output is — the average, produced from an under-specified request.

Generic output is usually not a model problem. It is what an under-specified request looks like when something has to fill the gaps.

An interview step forces the specification to exist before the generation starts. You can do this with no skill at all, by pasting one sentence in front of your request: "Before you build anything, ask me the eight questions whose answers would most change what you build." It costs nothing and it fixes most of what people install skills to fix.

Cost-saving mode: use the expensive model only where it shows.

The setup offers a choice between full power — the best model at every step — and a cost-saving mode that brings in Opus 4.8 and other models for parts of the process, keeping the expensive model for the parts that need it.

This is the most genuinely useful idea in the whole video, and it is presented as a footnote in a setup menu.

Here is why it matters. Building a website with an agent is not one task. It is a dozen, and they are not equally hard:

StepNeeds the best model?
Deciding the visual direction and layoutYes. This is taste and judgement.
Writing the headline and the positioningYes. Hardest text on the page.
Scaffolding the project, config, boilerplateNo. Any competent model does this.
Converting a chosen design into markupMostly no. It is transcription once decided.
Filling in placeholder copy and alt textNo.
Repetitive edits across many filesNo, and this is the highest-volume step.

Fable 5 costs $10 per million input tokens and $50 per million output. Output is the expensive leg, and generating a site is an output-heavy job. Paying that rate to write boilerplate is the actual waste — not that you used a premium model, but that you used it on the steps where nobody could tell the difference.

Take this one away even if you take nothing else. It applies to every agent workflow you have, not just websites: work out which steps are judgement and which are transcription, and stop paying judgement rates for transcription.

A critical pass at the start, not at the end.

The third idea is an adversarial step early in the process — the system takes a deliberately critical view before the work begins rather than reviewing at the end.

The ordering is the whole point. A critique at the end can only catch what is already built, and by then the layout, the structure and the direction are all committed. Everything the review finds is either cosmetic or expensive.

Moving it to the front means it is arguing about the plan, when changing your mind is free. In practice that means asking, before generation: what is the most obvious version of this page, and what would make it look like every other AI-generated site? Then avoiding those specific things on purpose.

Again, you can do this without any skill installed. It is one extra turn.

Two variables, two states each. Four runs.

"Driver and vehicle both matter" is the most testable claim in this whole category, and almost nobody tests it — which is convenient, because the skill is the part being given away and the model is the part doing the heavy lifting.

No skillSkill installed
Weaker modelA — the floorB — can instructions replace ability?
Fable 5C — how much was the model all along?D — the configuration being sold

The comparison every video shows is A against D. That gap is guaranteed to be huge and it proves nothing, because two things changed at once.

The two that mean something:

  • C versus D — same model, skill on and off. This is the only measurement of what the skill is worth.
  • B versus C — if you could only have one, which would you take? This is the actual question, and nobody answers it.

My prediction, written here so it can be wrong: C beats B comfortably, and D beats C by a modest margin. The model does most of the work; the skill is a real but smaller effect, and most of its contribution comes from the three process mechanisms above rather than from anything about design.

Anthropic's own framing of Fable 5 points the same way — they describe it as understanding intent with much less correcting and nudging, and catching design problems earlier models missed. That is ability living in the model, not in a document.

Get one number nobody has to agree with, and one they can.

These comparisons never settle anything because they are judged by the person who made the video, on how nice it looks.

Split the scoring in two.

The half that is not an opinion. Run every output through a checker that reads what the browser actually rendered and reports problems by category — text too small to read comfortably, contrast below the threshold, lines too long, headings out of order. These are measurable, they are where one-shot output reliably falls down, and nobody has to agree with the result. You get a count per cell.

The half that is. Score the look by hand, on a rubric you wrote before seeing any output, and judge with the labels stripped off so you do not know which cell you are looking at.

Report both. A page that wins on taste and loses badly on contrast is a normal outcome, and collapsing the two into one score hides exactly the thing you needed to see.

An independent benchmark page showing scored one-shot generation results for Claude Fable 5.
Someone else's one-shot scores. A useful sanity check, and no substitute for running your own brief.

The benchmarks and the professionals disagree, and both are right.

On paper the model is dominant. At launch Fable 5 ranked #1 on Code Arena: Frontend in every sub-category, and #1 on Design Arena for UI components, SVG and data visualisation with a 67% tournament win rate.

Then designers tried it on real work, and the reports are much colder.

  • Product leader Claire Vo asked it to one-shot a product design and described the result as grey, black and red with simple outlines — "not even AI-slop bad, fundamentally terrible design."
  • Designer Michal Malewicz tested it on real client projects and concluded that the impressive part ofs is precisely the part professional workflows do not need.
  • A comment on Hacker News landed on the most useful version: one-shot UI is "still ehhh", but the model "can keep a design system together much better without veering off into random tailwind classes."

These are not contradictory. They are measuring different things, and separating them tells you what to actually use it for.

What it is good atWhat it is not
Holding a design system consistently across a codebaseInventing an original visual direction from nothing
Producing a competent, conventional first draft fastProducing something a client will find distinctive
Implementing a direction you have already chosenChoosing the direction
Not drifting into random utility classes over a long sessionTaste, in the sense a designer means it

That left column is genuinely valuable and it is unglamorous, which is why nobody makes videos about it. Consistency across a hundred components is real work that models used to be bad at. Enforcement is where the model earns its money. Origination is where it does not.

It also explains the "$10,000 website" framing precisely. What an agency charges five figures for is mostly the origination — the discovery, the direction, the argument about what the thing should be. That is the parts skip and the designers are complaining about. The part being automated well is the part that came after.

It is very good at applying a decision and weak at making one. The expensive half of design work is making it.

The gap between a good first draft and a site you can ship.

One-shot generation has got genuinely good, ands are not faked. What follows is not a complaint about quality — it is a list of the things that are simply not in scope for a single generation, and which make up most of the remaining work:

  • Real content. Placeholder copy that reads plausibly is the single biggest difference between a demo and a live site, and writing the real version is the slow part.
  • Responsive behaviour at real breakpoints. Demos are shown at one width.
  • Accessibility beyond the automatic checks. Keyboard navigation, focus order, screen-reader labelling.
  • Performance. Generated pages tend to be image-heavy and unoptimised.
  • Anything with state. Forms that submit somewhere, authentication, a CMS someone else can edit.
  • Cross-browser checking. One screenshot in one browser is not testing.

None of this makes the tool less impressive. It does mean "one-shot" measures the first draft, which is a real and limited thing to measure. Judge it on that and it looks great; judge it on shipping and you will be disappointed for reasons that are not the model's fault.

What it costs, and what happened in June.

Price. Fable 5 is $10 per million input tokens and $50 per million output. Regenerating a whole site twenty times while you iterate is not a rounding error, and it is the reason the cost-saving mode above is the most practical idea in the video.

Availability. Fable 5 launched on 9 June 2026. On 12 June a US export-control directive suspended it for all users. The controls lifted on 30 June and access came back on 1 July. Nineteen days, on Anthropic's own record.

If your workflow is "I build client sites with this model", that is a business continuity fact rather than trivia. The mitigation is dull and effective: keep the brief and the skill portable, so the driver can be pointed at a different vehicle. Since a skill is a markdown file, this is easy — which is a genuine argument for using one.

There is also a smaller everyday version. Fable 5 routes a narrow band of sensitive requests to Claude Opus 4.8 rather than answering directly, which Anthropic says happens in under 5% of sessions. Most people building marketing sites will never notice.

Honest limits.

  • We have not run the four cells. This page publishes a test design and a stated prediction, not results. Do not cite a number from here that does not have a vendor page behind it.
  • One brief is one brief. A skill tuned for marketing pages may do nothing for a dashboard. Any result is about your brief, not about websites in general.
  • One run per cell is noise. Model output varies enough that you need at least three per cell, and the spread will be wide.
  • The $10,000 figure is an agency rate, not a valuation of generated output.
  • A free skill is still a supply chain. You are adding a document your agent reads on every relevant turn. Read it first — it takes two minutes, and that is the whole advantage of the format.
  • Image generation is usually a separate paid service. The media step in these workflows typically points at a third-party generator, which is its own bill and often the sponsor of the video.
  • Prices and model names move. Both were read on 10 August 2026.

Take the free skill. Take the three ideas instead of the file.

The skill is worth installing. It is free, you can read it in two minutes, and it will improve your first draft.

But the reason it works is not that it knows about design. It works because it asks you questions before building, spends the expensive model only where the expense shows, and criticises the plan while criticism is still cheap. Those three ideas belong in every agent workflow you have, and you can use all of them today without installing anything.

Then run the four cells on your own brief before you believe anyone's ranking — including the prediction on this page.

Everything above, traceable to a primary source.

Keep exploring

Browse more long-form resources for building and working with AI.