Swatch

Swatch is a small set of tests that helps me understand how AI models behave.
Why I made this
New AI models appear often. Each model has different strengths, limits, and habits.
Benchmarks give useful data. Release notes list new functions. These sources do not show what a model feels like during real work.
As a product manager, I need practical knowledge. I need to know how a model writes, sees, designs, researches, and uses tools.
I also need to know how the model handles unclear requests. Small behavior differences can change a product experience.
Swatch does not select the best model. It helps me select a suitable model for a specific task.
What Swatch contains
Swatch contains repeatable experiments. Each experiment tests one or more model behaviors.
One experiment gives a model three photos. The model must use the photos to write a short road-trip story.
Another experiment gives a model one image. The model must describe the image and then make a similar image.
I can use the same experiment with different models. I can then compare the actual outputs.
How I use it
- Select an experiment. Select a task that tests the behavior that you want to examine.
- Use the same input. Give each model the same prompt, instructions, and source files.
- Keep the output. Save the complete response. Do not replace an old response with a new response.
- Write what you notice. Record good work, errors, assumptions, and unexpected behavior.
- Compare the work. Put the outputs next to each other. Do not reduce the result to one score.
What I want to learn
I use each experiment to examine a part of the model. I want to understand how the model sees, thinks, and acts.
Imagination
Does the model make original connections? Can it create clear work with a distinct style?
Perception
What does the model see? Does it understand details, space, and visual relationships?
Taste
Can the model make good design choices? Does it use hierarchy, composition, and restraint?
Reasoning
Can the model organize unclear information? Does it make sound assumptions and set useful priorities?
Research
Can the model find reliable information? Can it combine sources and show evidence for its claims?
Agency
Can the model make a plan, use tools, recover from errors, and respond to a changed request?
Experiment map
Each experiment gives me a different view of the model. Together, the experiments create a useful model profile.
| Experiment | Test | What I learn |
|---|---|---|
| Write a story from three photos | Use three unrelated photos in one road-trip story. | Imagination, visual grounding, narrative, style |
| Describe a detailed image | Describe all relevant parts of a dense photograph. | Perception, space, detail, hallucination |
| Recreate an image | Describe an image and then make a similar image. | Visual translation and image fidelity |
| Improve a rough visual | Turn a rough visual and a brief into polished work. | Taste, composition, instruction use |
| Design a product homepage | Design a homepage from a short product description. | UI taste, hierarchy, copy, frontend skill |
| Build a page from a screenshot | Build a working page from one screenshot. | Perception, frontend skill, visual fidelity |
| Improve a weak interface | Improve an interface that has clear design problems. | Design judgment and restraint |
| Turn rough notes into a PRD | Turn incomplete meeting notes into a useful PRD. | Synthesis, ambiguity, product reasoning |
| Review a product | Review a product screen and recommend changes. | Product judgment, reasoning, priorities |
| Handle an unclear request | Respond to a broad request with limited direction. | Questions, assumptions, initiative |
| Research a question | Answer one question with eight to ten sources. | Search, synthesis, citation quality |
| Find facts in long documents | Find related facts in several long documents. | Retrieval, context use, attention |
| Analyze a messy spreadsheet | Use a difficult spreadsheet to answer a business question. | Analysis, tool use, quantitative reasoning |
| Complete a task with tools | Find options, compare them, and make an artifact. | Planning, tool selection, recovery |
| Adapt to a changed brief | Change an important requirement during the task. | Adaptability and context updates |
Swatch