e2e: end-to-end tests where you decide how much AI goes in
The idea in one sentence
You don’t have to choose between deterministic tests and AI tests: you choose how much AI each step carries, and you pay for the model once.
That is the pitch of e2e, by TesterArmy: an end-to-end testing framework for web and mobile (TypeScript, Apache-2.0) where you describe a goal in natural language and an agent drives the app to reach it, then check the result with locators and assertions in the same test. Everything below comes from its README and its documentation; I haven’t tried it.
Three kinds of step, one test
An e2e test combines three kinds of step, and only two of them call the model:
| Step | API | Calls the model? |
|---|---|---|
| Goal | agent.act |
Yes, unless the cache replays it |
| Assertion | agent.assert, agent.waitFor, agent.extract |
Yes |
| Locator | screen, expect |
No |
The example from its README:
test('a member upgrades to Pro', async ({ app, agent, screen }) => { await app.open('/settings/billing');
await agent.act('upgrade the workspace to the Pro plan'); await agent.assert('the invoice preview shows a prorated amount');
await expect(screen.getByRole('status')).toContainText('Pro');});Concept: the first line is an agent; the last one is plain old testing, deterministic and free. AI only goes where you put it.
Actions:
- Use
agent.actfor flows (what has to be achieved) and a locator for the exact result (what must be on screen). - Locators in this order of preference: role and name, then label, and test id last.
- For values that change on every run (an email with a timestamp), use
unique(), so the step can be replayed.
The replay cache: pay for the model once
This is what solves the economics of AI testing. An agent step that a later check verifies gets recorded, and the next run replays it without calling the model until the app changes. Each run’s summary reports it:
AI 4.1k tokens · 2 model calls · anthropic/claude-sonnet-4.5Cache 4 replayed · 1 handed off · 1 missedIt is not “record and pray”. On replay, the runner:
- Finds each recorded control by role, name, test id and context, and stops if the match is ambiguous.
- Checks the final route and what the step changed: which controls appeared (with their state: checked, expanded…) and which went away.
- Requires at least one change to happen during the replay: an outcome already on screen before the actions that should produce it proves nothing.
- If something doesn’t add up, it doesn’t fail: it degrades. The agent carries on from the current screen with a record of what ran and why replay stopped (the
handed offin the summary).
Details that show they’ve thought it through:
- Ids, tokens and timestamps in the URL are read as a placeholder:
/orders/42?t=1727780000and/orders/7?t=1727780999are the same route.?mode=safeand?mode=unsafeare not. - Text that changes on every run (dates, durations, ids) is not used for comparison.
- A recording is only saved if a later verification passes: a locator matcher,
agent.assertoragent.waitFor. Anexpect(value).toBe()on a plain value or anagent.extractdon’t count.
Action: always put a check of the outcome right after each agent.act. Without one the step is never recorded, and you pay for the model on every CI run.
Failing is not the same as not knowing
agent.assert separates two outcomes that almost no tool tells apart:
ASSERTION_FAILED: the statement is false.ASSERTION_INCONCLUSIVE: there is not enough evidence on screen (the value is on another page, still loading, or only in the pixels).
A test that can’t see the outcome shouldn’t claim the app is broken.
The other two assertions:
agent.waitForwaits for a condition and only calls the model when the screen changes.agent.extractreads data from the screen into a type (Zod, or any Standard Schema), which you then check with a regularexpect.
With vision: true the judgment includes a masked screenshot. They warn that images add tokens, and after a secret is filled no more screenshots are sent for that step.
Decision models instead of a chat
The @e2e-dev/decision package runs agent.act and agent.assert through a decision model instead of an LLM. For each action it answers a single question: which operation, and on which element, choosing among the options built from the screen’s semantic tree. It never generates free text.
- When the operation is typing, a small text model writes the value. Model output never becomes a selector, a URL or a key, and password fields never reach it.
- Any provider that answers choice questions with probability distributions works; their example uses
typeSafeAi.decisionModel('jev-latest'). - Two optional thresholds,
minProbabilityandminConfidence, block doubtful actions and turn weak assertions into inconclusive ones. They point out that a threshold is your policy choice: measure it on your own tests before changing it.
It is the idea behind Jev applied to a concrete case: many small, typed decisions with a probability attached, instead of a model that converses.
Bug bashes: only bugs with a failing test
The most ambitious part. A bug bash runs several exploration sessions at once, each on one area of the app, and only reports as a bug what a repro test confirms.
Concept:
- The agent reads the app’s routes and, on a branch, the diff.
- It writes 5 to 10 charters: a one-sentence goal for one area, with one posture (first-time user, numbers and copy, edge input, state after reload, error paths). Each runs with its own persona, four at a time.
- It rejects findings with a known cause: the local environment, the intended design, the seed data, or the explorer’s own limits (a link that opens a new tab, for example).
- It writes a repro test per finding, asserting the expected behaviour. The finding is a bug only if that test fails.
Actions:
- Repro tests are tagged
bugbash: keep them out of the run that gates merges withe2e run --exclude-tag bugbash. - Once the bug is fixed the test passes: remove the tag and keep it as a regression test.
- The explorer’s video is not masked: if it typed a password, check the video before sharing it.
Built for coding agents
- An MCP server with live sessions, several at once, so an agent can check locators against the app.
- A skill (
npx skills add tester-army/e2e) with topics such asbug-bash. - The whole documentation ships inside the npm package, in
node_modules/e2e/docs, so agents can read it offline.
Caveats
- Very young (created in July 2026) and on the way to 1.0: the authors warn that APIs and config can still change between minor releases.
- It does not sandbox test code: it runs with your permissions. Untrusted PRs need an external sandbox without secrets or write tokens.
- It sends anonymous telemetry by default (commands, engines, where runs fail; no content or credentials). Turn it off with
E2E_TELEMETRY_DISABLED=1. - Platforms: web through Playwright (Chromium, Firefox, WebKit) and mobile through agent-device (iOS simulators and Android emulators), with the same API. Platform-agnostic, not language-agnostic: it’s TypeScript.
Where to start
There are migration guides from Playwright, Cypress, Selenium, Detox and Maestro, all incremental:
- The Playwright one suggests running e2e alongside Playwright and migrating one file at a time.
- The Cypress one keeps your
data-cyattributes (web({ testIdAttribute: 'data-cy' })).
The sensible move: pick a flow that is fragile today (lots of selectors, changes often), rewrite it with an agent.act followed by a deterministic check, and watch in the CI summary how many steps replay for free and how many end up with the agent.
My take
Three things stay with me:
- How much AI is decided step by step. That’s the right way to frame it: the agent to get there, the locator to check. There’s no need to bet the whole test on a model.
- The replay cache is what makes it viable in CI. Paying for the model only when the app changes, and falling back to the agent instead of failing, turns the cost of every run into the cost of every change. And the rule that an effect already on screen proves nothing is exactly the kind of detail that separates a serious tool from a demo.
INCONCLUSIVEand bugs with a repro are honesty in practice. Neither the test nor the bug bash reports what it can’t prove. It’s what I’d most like to see in any tool that uses a model to judge.
I’d try it on one fragile flow, keeping an eye on telemetry and isolation, but I wouldn’t yet make it the base of a product’s whole suite: 1.0 hasn’t arrived.
Bibliography:
- tester-army/e2e: Next generation e2e testing framework for web and mobile apps. TesterArmy
- Core concepts. e2e documentation
- Caching. e2e documentation
- Decision models. e2e documentation
- Bug bashes. e2e documentation
- Security. e2e documentation