Evaluation-and-methodology.md

Evaluation plan and methodology

What this report did

Reviewed primary vendor documentation and product pages, an open model repository, and Kuaishou's reported financial results. Research was performed September 26, 2026. Sources were checked for current availability where possible; old announcements were not automatically treated as the latest release.

The report separates three kinds of statements: documented offerings, vendor-reported performance or business results, and our analysis. It does not reproduce vendor benchmark wins as an independent ranking. It includes no private company wiki, diary, acquisition discussion, customer data or internal product strategy.

No paid model calls, customer interviews, reproducible output comparisons or revenue cohort analyses were conducted. Consequently, this is a sourced market and strategy brief with a concrete next research plan, not a completed empirical benchmark or investment recommendation.

A practical evaluation set

Build a consented, reusable set of briefs across the intended customer's jobs. Include product fidelity, character continuity, controlled camera motion, dialogue, localization, editing an existing shot and a short multi-shot sequence. Use both straightforward requests and the cases the customer currently rejects.

For each brief, define the non-negotiables before generating: exact product details, required words, permissible variations, reference permissions, duration and delivery format. Give every provider the information and controls it actually supports; record any manual intervention.

Use multiple generations per brief and preserve failed outputs. Hide provider names during human scoring where feasible. Score at least:

The sample size should be chosen for the decision being made. A small pilot can identify obvious incompatibilities; it cannot establish a universal leaderboard. Report uncertainty and per-task outcomes rather than averaging unlike jobs into one score.

Measure the whole system

Separate model time, queue time, human review and editing. Save the exact model/version, endpoint, settings, references and date. When a supplier changes the model behind a stable endpoint, rerun a fixed set of cases to detect improvement or regression.

Test interrupted jobs, retries, quota exhaustion and unavailable providers. Check that references and output assets can be exported with their metadata. An appealing demo is less useful if a routine interruption loses the approved work.

Questions still open

Which buyer will pay for the complete workflow rather than use an existing creative suite? How much retry and editing cost is hidden by curated demos? Does a team's accumulated asset history create meaningful switching cost? Which controls remain reliable after a model update? What share of usage is recurring production versus one-off experimentation?

Private-company revenue and profitability are particularly uncertain. Funding announcements, registered users, generated-video counts and annualized run rates should not be treated as interchangeable with recognized revenue or retention. This report deliberately does not estimate a total market size from those proxies.

Refresh procedure

Before a buying or integration decision, recheck the provider's current availability, API documentation, pricing and applicable product terms. Update the source ledger with the observation date and record what changed. Keep the earlier edition so readers can distinguish a historical observation from a current claim.

Back to the report

Open with JavaScript for the full viewer.