Methodology
How the board is scored
Ten axes, weights summing to 100, four tier bands and a community score from 0 to 1000. Everything on this page is the actual arithmetic behind the ranking, including the part where we tell you which votes are ours.
Launch scores include seeded votes from our own testing panel
A scoreboard that opens with zero votes is not a scoreboard, so before launch our own testing panel ran the same brief on every tool on the board and voted on what it found. Those votes are part of the current scores. They are badged Editor panel on every opinion card and in the public data, so you can always tell them apart from community submissions, and we do not blend them into an anonymous average.
Community votes accumulate on top of the seed and are never adjusted to protect it. As volume grows the panel contribution shrinks as a share of the total, which is the intended direction. If a community consensus contradicts a panel score, the community consensus wins, because that is what the board is for.
- 109
- Panel opinions
- 0
- Community opinions
- 76,242
- Axis votes counted
- 10
- Axes per tool
The weights
Ten axes, summing to 100
| Axis | Weight | Share | What the vote is actually measuring |
|---|---|---|---|
| Design | 22 | Does the first render look like something you would show a client, before any cleanup pass. | |
| Speed | 21 | Cold start to first interactive render, and how long a normal iteration takes after that. | |
| Agent performance | 11 | How far it gets on a multi step brief with no human intervening, and how it recovers from its own mistakes. | |
| Reliability | 9 | Does the project survive week two. Do schema changes propagate instead of quietly breaking things. | |
| Integrations | 9 | Auth, database, storage, email and payments. Wired already, or a service account and an afternoon. | |
| SEO and GEO | 8 | Server rendered output, metadata, structured data, and whether an LLM can read the result. | |
| Scalability | 6 | What happens at a few dozen files, a few thousand records and a second developer. | |
| Value | 6 | What a real working month costs, not the sticker price of the entry plan. | |
| API and MCP | 5 | Programmatic access, and whether you can point your own agent at it. | |
| Code ownership | 3 | Can you take the code and leave. Full export, or a hosted runtime you cannot reproduce. | |
| Total | 100 |
This board weights design and raw speed far more heavily than a typical enterprise scorecard
Design carries 22 points and speed carries 21. Together that is 43 of the 100, nearly half the score decided by how good the output looks and how fast you get there. An enterprise scorecard typically allocates 10 to 20 points across those two axes and spends the rest on governance, compliance, support terms and vendor viability.
We weight it this way because it reflects what the building community actually votes on. When we ask people what decides which tool they open, they answer with the first render and the wait. Nobody starting a project on a Tuesday afternoon opens a vendor risk assessment first.
That is why our ranking differs from other boards, and that is the point. If your constraints are procurement constraints, take the public data, apply your own weights, and you will get a different and equally valid order. We would rather publish a board with a stated bias than one that pretends it has none.
Tier bands
How a band is earned
S tier
8 entries
Completes the full board brief, a real application with data, auth and a deploy target, and the output is presentable at the end.
A tier
10 entries
Completes the brief with a compromise voters can name: serviceable design, unpredictable cost, or a plan that fails late.
B tier
20 entries
Completes a narrower version of the brief, or completes it in a way that does not survive several rounds of iteration.
C tier
13 entries
Does not attempt the brief. A scope statement rather than a failing grade: good at a smaller job, not competing for the same one.
The brief every tool runs
Every entry is asked to build the same thing: a small multi user application with a data model of at least three related tables, sign in, a list view with filtering, a detail view with editing, file upload and a working deploy. Same prompt sequence, same order, no hand written code and no rescuing a tool that gets stuck. Where a tool cannot do part of the brief, that is recorded rather than worked around.
From votes to a score out of 1000
Each axis collects its own votes. An axis score is the share of positive votes on that axis, and the community score is the weighted sum of the ten axis scores scaled to 1000. That is the whole calculation. There is no editorial adjustment layer, no vendor input and no normalisation curve applied after the fact.
We publish the raw vote count next to every score for a specific reason: a weighted average with no denominator is close to useless. Two tools on 900 are making very different claims if one has four thousand votes and the other has forty, and a star rating hides exactly that difference. This is also why there are no stars anywhere on the board and no 0 to 100 index.
Where a 0 to 5 rating appears
One place only. Structured data consumers expect a conventional rating, so tool pages emit a machine readable value derived honestly from the community score by dividing it by 200, with the vote count as the rating count. It is a translation for machines, not a second scoring system, and it never appears in the interface.
Duels and the head to head record
Alongside axis voting we run duels: two tools, one unweighted question, which would you start with. The W and L record on every row is the sum of those duels. Duel results and board scores disagree regularly, because an aggregate and a forced choice measure different things. We publish both rather than reconciling them into one number that would hide the disagreement.
Weekly rank freeze
Scores move continuously. Board rank is snapshotted every Monday and held for the week so that trend arrows mean something and so the board can be cited at a stable position. Every change is listed on the movers page with a written reason drawn from the vote data.
Corrections
Each tool page carries a Last verified date from its own row, and the board carries a Board updated stamp. Neither is hardcoded. When a row turns out to be wrong we change it, update the verified date, and publish the correction as a dated Pulse report rather than editing quietly. If you spot something wrong, the contact page routes corrections straight to us.
Conflicts of interest
Agent Verdict is an independent community scoreboard and is not operated by any vendor listed on it. No entry pays to be listed, no entry pays for placement, and no vendor sees a score before it publishes. Links to vendor sites carry nofollow. If we ever take money from anyone on the board, that will be disclosed on this page before it happens rather than after.
Questions
About the method
Why does design carry 22 points?
Because it is the axis that decides whether the output of an AI app builder is usable at all. A generated application that works but looks generated still needs a designer before anyone will show it to a customer, and that is the single largest hidden cost in this category. Voters consistently rank it first when asked what they actually judge on.
Why is this board weighted so differently from enterprise scorecards?
An enterprise scorecard is written for someone signing a multi year contract, so it spends most of its points on governance, support terms and vendor stability. This board is written for someone starting a project this week. Both are legitimate; ours is explicit about which question it answers, and the weights are published so you can recompute the ranking with your own.
Are the launch scores real votes?
They are real votes from a real testing round, cast by our own panel and badged Editor panel everywhere they appear. They are not fabricated numbers and they are not vendor supplied. Community votes accumulate on top and will outweigh the seed as volume grows.
Can I get the underlying data?
Yes. The whole board is published as JSON and CSV under a CC-BY 4.0 licence, including tier, score, vote counts, trend and head to head record. Cite Agent Verdict and keep the vote count next to any score you quote.
How do you stop vendors gaming the vote?
One vote per axis per browser, an anonymous first party cookie rather than an account, and manual review on every written opinion before it publishes. It is not unbeatable and we do not claim it is. What makes the board hard to game is that the vote counts are published, so an implausible spike is visible to everyone reading it.