1. Separate the deliverables
A product photo, a static campaign and a generated clip need different acceptance conditions. Pomelli is listed for its observed static campaign workflow; Lovart for documented design-agent canvas work; Kling for documented video and Canvas Agent features. The remaining tools provide different image or creative routes. A source-page check does not measure the finished asset.
2. Freeze the input before generating
Use an approved SKU fact sheet and licensed reference images for a commercial comparison. Record name, material, dimensions, quantity, audience, claims allowed, exact copy, aspect ratio and expected asset count. Label synthetic fixtures as synthetic. Lock the checks before sending the request; keep the first output even when the result disappoints.
3. Include the existing brand context
Record the workspace business profile and any automatically selected image ingredients. A text brief can be affected by existing context. The retained Pomelli test uses an existing RoboSkin profile with an unrelated fictional CedarDesk brief; its result only assesses that configuration. New-brand onboarding and removal of the existing ingredient would be separate tests rather than repairs of this first attempt.
4. Check the artifact, not just the brief
Inspect every output for product shape, text, colour, material and invented claims. Measure actual exported dimensions. For video, inspect continuity, duration, aspect ratio, captions, voice and all scenes. Retain failures and editing time. Account access and a plausible script do not establish a usable video.
5. Compare costs only when measured
Include plan fees, actual credits, all attempts, rejects and review time. No checkout is not a measurement of total cost. A model label on a marketing page is not an observed backend identity. Publish only the scope that was tested, and link to the complete evidence.