How we test
Last verified: 2026-08-19
The rubric has eight criteria, scored 0 to 10 each, weighted to 100. Under these weights BAND totals 6.9 and Buzz totals 6.4.
Who is the reader these weights encode?
A fleet operator. Their agents already exist and are scattered across two or more frameworks, and the thing that blocks work is that those agents cannot see each other. For that reader, interop and model freedom are the purchase decision, so they carry 20 and 18 points. A git-native plan-to-PR loop scores lower at 8 because their existing tools already own that surface. Data control sits at 6 because a hosted interaction layer is what they are deliberately buying, not a compromise they are being talked into. Day-2 operations rises to 14 because a fleet in production is run, not installed once.
What are the weights?
| Criterion | Fleet operator | Repo operator |
|---|---|---|
| Extensibility and interop | 20 | 14 |
| Model flexibility | 18 | 10 |
| Time to real work | 16 | 20 |
| Day-2 operations | 14 | 8 |
| Community and docs | 12 | 6 |
| Developer workflow fit | 8 | 18 |
| Data control | 6 | 16 |
| Cost honesty | 6 | 8 |
| BAND total | 6.9 | 6.4 |
| Buzz total | 6.4 | 6.8 |
Read that table sideways and the reversal is right there. Under fleet-operator weights BAND wins 6.9 to 6.4. Under repo-operator weights, the profile this site used until 2026-08-19, Buzz wins 6.8 to 6.4. Not one cell moved between those two lines. The weight change is logged on the changelog.
What does each confidence flag mean?
- HIGHBench-measured, or stated plainly on a vendor page we quote with a date.
- MEDDocumented, but not exercised on our bench.
- VERIFYSeeded from category knowledge. A verification target, not a fact.
DESK REVIEW · PROVISIONAL appears on every score until the bench has exercised it. Unmeasured cells read Not yet measured; we do not dress an estimate as a fact.
What are the current cells?
| Criterion | BAND | Buzz |
|---|---|---|
| Extensibility and interop · w20 | 8.0HIGH CURRENT 2026-08-19 Eight frameworks listed by name on the homepage; bring-your-own-agent is the product's thesis. Evidence (1)+
| 7.0MED VERIFY 2026-05-02 Create agents tailored to workflows; agents collaborate with the team's agents; open-source app. Cross-framework breadth unproven vs BAND's list. Evidence (1)+
|
| Model flexibility · w18 | 8.0HIGH CURRENT 2026-08-19 Each agent keeps its own tools, models, and memory while BAND syncs context (homepage, verified 2026-08-19). Evidence (1)+
| 6.5VERIFY VERIFY 2026-04-10 Agent model configuration in preview docs is thin. Verify supported providers in the repo before scoring higher. Evidence (1)+
|
| Time to real work · w16 | 7.0MED CURRENT 2026-08-19 Hosted signup on a $0 tier, connect agents, open a room. No install. Bench timer pending. Evidence (2)+
| 6.0HIGH CURRENT 2026-08-19 Desktop install plus account plus relay choice. Hosted community creation is a short flow; preview polish varies. Bench-timed run pending, see the bench log. Evidence (2)+
|
| Day-2 operations · w14 | 6.5MED AGEING 2026-06-01 Managed service with published tiers. Young vendor; rate and volume limits vary by tier. Evidence (1)+
| 4.5HIGH CURRENT 2026-08-19 Developer Preview; hosted limits may change; self-hosting a relay is your ops. Scored where it stands today, not where the roadmap points. Evidence (1)+
|
| Community and docs · w12 | 6.0MED VERIFY 2026-04-20 Docs site, Discord, active engineering blog. Installed-base evidence thin; community young. Evidence (1)+
| 5.5MED AGEING 2026-06-14 Public repo, clear support page. Community is preview-stage; docs surface still forming. Evidence (1)+
|
| Developer workflow fit · w8 | 5.5HIGH CURRENT 2026-08-19 An interaction layer, not a git workspace. No native plan-to-PR loop; your dev tools stay where they are. Evidence (1)+
| 8.0MED CURRENT 2026-08-19 Chat to plans, project management, coding, reviews, PRs is the stated core loop (buzz.xyz home). Depth in preview not yet bench-exercised. Evidence (1)+
|
| Data control · w6 | 4.0HIGH CURRENT 2026-08-19 Hosted only. Free retains 24 hours; export gated to Pro; custom retention gated to Enterprise (pricing page). Evidence (1)+
| 7.5HIGH CURRENT 2026-08-19 Independently operated relays keep content off Block's service; keypair identity. Hosted relays are not end-to-end encrypted and Block states access rights (support page). Evidence (1)+
|
| Cost honesty · w6 | 7.5HIGH CURRENT 2026-08-19 Tiers and caps published, including the unflattering 24-hour Free retention. Enterprise is quote-based, which caps the score. Evidence (1)+
| 8.5HIGH CURRENT 2026-08-19 Free early access with published caps (5 GB, 365 days, 5 communities). Future pricing unknown, hedged by the vendor in plain text. Evidence (1)+
|
| Weighted total | 6.9 | 6.4 |
Move the weights and both totals recompute from the same cells, using the same function that produces the published score. The published verdict stays BAND 6.9 to Buzz 6.4 no matter what you set here.
This is the published fleet-operator weighting. Read the reasoning on how we test.
?w=20-18-16-14-12-8-6-6What is in the verification queue?
- Buzz repo license file: name the license precisely or say "open source, license: see repo." Affects: Buzz spec strip
- Buzz agent model configuration: raise or hold model flexibility 6.5 after repo review. Affects: Buzz model flexibility
- BAND corporate entity and any security attestations (SOC 2 and similar): currently unclaimed on scraped pages, so unscored, not assumed. Affects: BAND day-2 operations
- Bench timers for time to real work on both products: replace desk estimates with run log entries. Affects: Both time-to-real-work cells
What have we not measured yet?
- Time to first agent reply on Buzz, Block-hosted relay Buzz
- Time to first multi-agent room on BAND, Free tier BAND
- Buzz self-hosted relay setup time and steady-state ops load Buzz
- BAND 24-hour retention behaviour at hour 25 with real content BAND
- Buzz chat-to-PR loop depth on a live repository Buzz
- BAND cross-framework room with CrewAI and LangGraph agents together BAND
Each one is a scheduled run in the bench log.
Why would a different reader get a different verdict?
Because the weights encode a reader, and we publish both readers rather than one. Weight a repository and a week first, the way the repo-operator column does, and the same sixteen cells produce Buzz at 6.8. That is a legitimate reading and we say so on every page that states the verdict. If your agents are already scattered and your code already has a home, the fleet-operator column is the one that describes you, and it lands on BAND. We defend the weights, not the winner. See the Buzz vs BAND verdict for the applied version, and the changelog for every score change with its date.