Don’t ask only, “Which model made this?” Ask, “What would I need in order to use, change and trust this?”
This is our original evaluation framework. It is deliberately independent of a model ranking, and it does not claim we have repeated every demo in the showcase.
01 / Separate the result from the claim
A video proves that an attractive sequence was recorded. It does not reveal how many attempts preceded it, how much a person corrected, or whether the file can be edited. Before copying a workflow, write down the output you actually need: a screenshot, a navigable website, an editable scene, or a deployable application. These are different deliveries. If your client needs to change a product colour next week, a beautiful flattened image is not a substitute for the editable scene.
02 / Ask for the smallest inspectable artifact
For a website, open the live page and try its main action on a phone. For a 3D scene, look for a native project or an export that can be opened independently. For an automation, inspect a complete input and output, including a failed run. You do not need the author’s entire environment to learn something; you do need enough evidence to distinguish a finished output from an illustration. When that evidence is missing, keep the case as inspiration rather than a proven workflow.
03 / Name the whole stack
Record the product, model label, reasoning setting, connected tools and imported assets separately. “Made with Claude” could describe an assistant writing Python for Blender, a model directing another service, or a human following suggested steps. The same final render can conceal very different dependencies. If a model version is not stated, write “not disclosed”. Do not fill it in from the post’s date. Likewise, a ChatGPT image does not establish the image-generation ability of the text model selected in the conversation.
04 / Run a change request, not just a replay
If you choose to test a workflow, change one meaningful requirement. For a landing page, replace the headline with a longer one and switch to a narrow screen. For a product scene, change the material and camera angle. For an analysis, change an input and check whether the conclusion changes for the right reason. This is our suggested evaluation method, not a claim that we have run these tests on the examples in this preview. A workflow that only reproduces a single polished demo may still be useful, but its limits should be visible.
05 / Make the cost comparable
Keep subscription fees, per-token API prices and the cost of completing a task in separate columns. Include the attempts you discarded, third-party assets and the time spent correcting the output. A subscription you already have may be the best starting point even when another model leads a benchmark. Set a stop rule before experimenting: one small deliverable, a time limit you choose, and a maximum additional spend. This prevents an appealing demo from turning into an open-ended tool-shopping exercise.
06 / Decide what the evidence earns
Use three practical conclusions: “inspiration”, “worth a bounded test”, or “ready for my use”. Inspiration needs a visible result and a credible source. A bounded test additionally needs a known path to reproduce or adapt it. Ready-for-use requires your own checks on the actual output, rights, privacy and failure behaviour. None of these labels means a model is universally best. Keeping the conclusion narrow is what makes a collection useful instead of another leaderboard.