EmailBench scoring rubric
TL;DR: Every tool on EmailBench is scored 0–10 on eight weighted criteria. Overall scores are weighted averages. Category rankings use the matching criterion (or a published composite). We do not invent funding figures, customer counts, or fake quotes.
Why we publish the rubric
Review sites that hide methodology produce vibes, not decisions. EmailBench exists so a lifecycle marketer can see why a tool scored 9.2 on automations and 7.4 on data — and decide which number matters for their stack.
Criteria and weights
| Criterion | Weight | Definition |
|---|---|---|
| Sending | 15% | Native delivery quality signals, authentication tooling (SPF/DKIM/DMARC), reputation controls, and reliability of the send path. |
| Design | 15% | On-brand visual quality, cross-client rendering, brand extraction or design systems, and speed from idea to polished HTML. |
| Automations | 15% | Lifecycle flows, triggers, branching, prompt-to-automation, and how quickly a team can ship welcome/cart/winback programs. |
| AI Generation | 15% | Natural-language campaign creation, brand-aware copy and layout, agent operability, and consistency across sends. |
| Data & Segmentation | 10% | Profile depth, event ingestion, predictive models, ecommerce/product graphs, and segment flexibility. |
| Developer & Agents | 15% | REST API quality, SDKs, MCP/agent interfaces, webhooks, and whether agents can create, schedule, and send without a dashboard. |
| Ease of Use | 8% | Time-to-first-send, learning curve for marketers, clarity of UX, and how much specialist help is required. |
| Value | 7% | Pricing clarity, free/trial posture, and whether capabilities match cost for the intended buyer. |
How overall scores are calculated
Overall = Σ (criterion score × weight). Displayed overall scores are rounded to one decimal. The AggregateRating values in our JSON-LD match the numbers shown on each review page.
Category rankings
Sending, Design, and Automations sort primarily on those criteria. Ecommerce uses a published composite (data, automations, sending, design). AI-native privileges AI generation and developer/agent surface.
What counts as evidence
- Hands-on product evaluation against each criterion
- Vendor documentation, pricing pages, and public product surfaces
- Public accolades only as adoption signals (e.g. Product Hunt wins)
- Revision notes when scores change (see review “Updated” dates)
What we refuse to invent
Specific statistics we cannot verify, funding rounds, customer counts, or quotes attributed to real people or companies. Qualitative comparison is fine; fabricated precision is not.
Affiliate disclosure
Some outbound links may be affiliate links. Rankings and scores are not for sale. Compensation never buys a criterion point. Full policy on the About page.
Revision history
- 2025-11-01 — Rubric v1 published with eight criteria.
- 2026-06 — Added explicit Developer & Agents criterion weight for MCP/API evaluation; rebalanced Ease/Value slightly.
- 2026-08-01 — Category composites documented for ecommerce and AI-native boards; score refresh across the desk.