# Amazap "Page to Markdown" v4 (after the post-v3 fix pass): final benchmark

*Run: Saturday 2026-10-10. **First buy issued at ≈12:51:05 CEST.** The first receipt (NPR, `amazap_mcp_mv29xcji_wjly86`) has a `fetchedAt` of 12:51:10.8 CEST. The last buy finished at ≈12:55:30 CEST. Same 20 URLs as v1–v3, plus 6 "hard" pages. Raw HTML was re-fetched at 12:50:57 CEST (`html_v4/`, with the hard-page result codes in `html_v4/fetch.log`). Default parameters throughout: `mode=article`, `frontmatter=false`, no `max_chars`.*

## Spend and balance consistency

* **26 buy attempts that charged** (cap 27):
  * 20 benchmark pages
  * 6 hard pages: the 5 planned, plus AP News, which doubled as the quote-docs test
* **Gross charged: 338 sats** (budget 350). **Credited back: 65 sats** (5 site blocks). **Net spent: 273 sats** = 21 delivered × 13.
* **Uncharged calls:** one `quote_required` (deliberate test, see below).
* **Balance:** 5,772 before → **5,499** after. 5,772 − 273 = 5,499 ✓
  * `spentTodaySats` went from 818 to 1,091 (+273) ✓
  * `orders` lists exactly these 26 orders since v3, with no foreign activity.
* **Lost-update race check: fixed.**
  * The 3 simultaneous buys (NPR, MDN 402, Stripe charges) returned **distinct, sequential** `balanceAfterSats` of 5,759 / 5,746 / 5,733, with receipt strings "5772 → 5759", "5759 → 5746" and "5746 → 5733". In v3 all three said 5,993.
  * All 21 delivered receipts chain exactly, each 13 below the previous one.
  * The 5 credited (blocked) receipts carry `balanceSats: 5499` (unchanged) but **no `balanceAfterSats` field**. That is a minor inconsistency in the receipt shape.
  * The receipts' `remainingSessionSats` (727 after H2) disagrees with `balance` → `spentSessionSats: 0` / `remainingSessionSats: 1000`. That looks like two different "session" notions. It is cosmetic.
  * New limits are visible since v3: `maxCallSats` 500, `maxDaySats` 2,000, `maxSessionSats` 1,000, `autoApproveSats` 50.

| # | Page | Order id | Status | Charged | Credited | Receipt balanceAfter | Expected | OK |
|---|---|---|---|---:|---:|---:|---:|---|
| 1 | 13 | `amazap_mcp_mv29xcji_wjly86` | ok | 13 | 0 | 5759 | 5759 | ✓ |
| 2 | 04 | `amazap_mcp_mv29xd0v_6rnnvm` | ok | 13 | 0 | 5746 | 5746 | ✓ |
| 3 | 07 | `amazap_mcp_mv29xe2l_qk3snl` | ok | 13 | 0 | 5733 | 5733 | ✓ |
| 4 | 01 | `amazap_mcp_mv29xzb6_sw4ndq` | ok | 13 | 0 | 5720 | 5720 | ✓ |
| 5 | 02 | `amazap_mcp_mv29y4hi_h7m5o7` | ok | 13 | 0 | 5707 | 5707 | ✓ |
| 6 | 03 | `amazap_mcp_mv29yes1_j8k1si` | ok | 13 | 0 | 5694 | 5694 | ✓ |
| 7 | 05 | `amazap_mcp_mv29yidl_x39gdn` | ok | 13 | 0 | 5681 | 5681 | ✓ |
| 8 | 06 | `amazap_mcp_mv29ynyu_trs8sd` | ok | 13 | 0 | 5668 | 5668 | ✓ |
| 9 | 08 | `amazap_mcp_mv29ytc2_9kiuuf` | ok | 13 | 0 | 5655 | 5655 | ✓ |
| 10 | 09 | `amazap_mcp_mv29z972_ppcnla` | ok | 13 | 0 | 5642 | 5642 | ✓ |
| 11 | 10 | `amazap_mcp_mv29zecv_8p2yxz` | ok | 13 | 0 | 5629 | 5629 | ✓ |
| 12 | 11 | `amazap_mcp_mv29zlnz_82jdfd` | ok | 13 | 0 | 5616 | 5616 | ✓ |
| 13 | 12 | `amazap_mcp_mv29zpt6_ux787h` | ok | 13 | 0 | 5603 | 5603 | ✓ |
| 14 | 14 | `amazap_mcp_mv29zuho_pfb5w2` | ok | 13 | 0 | 5590 | 5590 | ✓ |
| 15 | 15 | `amazap_mcp_mv29zzd8_xz8gcg` | ok | 13 | 0 | 5577 | 5577 | ✓ |
| 16 | 16 | `amazap_mcp_mv2a0bl9_jejqqq` | ok | 13 | 0 | 5564 | 5564 | ✓ |
| 17 | 17 | `amazap_mcp_mv2a0fe7_w9tv4u` | ok | 13 | 0 | 5551 | 5551 | ✓ |
| 18 | 18 | `amazap_mcp_mv2a0kic_be1sj5` | ok | 13 | 0 | 5538 | 5538 | ✓ |
| 19 | 19 | `amazap_mcp_mv2a0p61_159o37` | ok | 13 | 0 | 5525 | 5525 | ✓ |
| 20 | 20 | `amazap_mcp_mv2a0t57_qlbu66` | ok | 13 | 0 | 5512 | 5512 | ✓ |
| 21 | H1 | `amazap_mcp_mv2a13vt_t8gjaf` | ok | 13 | 0 | 5499 | 5499 | ✓ |
| 22 | H2 | `amazap_mcp_mv2a18g9_k0kjt5` | delivered_incomplete | 13 | 13 | (balanceSats 5499) | 5499 | ✓ |
| 23 | H3 | `amazap_mcp_mv2a1e3p_b8761u` | delivered_incomplete | 13 | 13 | (balanceSats 5499) | 5499 | ✓ |
| 24 | H4 | `amazap_mcp_mv2a1gqd_ks8gye` | delivered_incomplete | 13 | 13 | (balanceSats 5499) | 5499 | ✓ |
| 25 | H5 | `amazap_mcp_mv2a1jdy_z2nmh7` | delivered_incomplete | 13 | 13 | (balanceSats 5499) | 5499 | ✓ |
| 26 | H6 | `amazap_mcp_mv2a31hc_yn190k` | delivered_incomplete | 13 | 13 | (balanceSats 5499) | 5499 | ✓ |

## TL;DR

* **The 20-page set is essentially unchanged from v3, and slightly leaner.**
  * Median **1,562 tokens/page** (v3: 1,614; trafilatura: 1,954; text-only: 3,436; raw HTML: 107,522). Total 57,646 (v3: 58,464).
  * Median clean recall is **94.4%** (v3: 94.1%).
  * Body-heading retention is a median **100%**, aggregate 95% of all main-content h2/h3 (trafilatura-md: 84%).
  * Code: 100 fences vs 102 `<pre>` in the sources. The drop of one vs v3 is the removed Stripe duplicate, so it is correct.
  * Junk share is a median 1.3% (aggregate 3.5%). Link/image-URL overhead is **0%** on all 20 pages.
* **Size vs the baselines:**
  * −98.4% vs raw HTML (median per page).
  * −49.5% vs text-only. It is smaller than text-only on 20/20 pages.
  * vs trafilatura: **+5.5% per page** (median ratio 1.055×; v3 1.08×), but **−16.6% in aggregate**. It is still larger on 17/20 pages, by a median of 55 tokens (v3: 88).
* **Most of the v3 "still broken" list is fixed.** Title leaks, empty links, indented-text code blocks, the Stripe duplicate fence and its missing headings, the NPR credit and the stray "Top" are all gone. **Two items are not fixed:**
  * Wikipedia math is still missing.
  * Nav/caption trimming is only marginally better.

  The `quote_required` fix is partial: the error is now actionable, but the docs are unchanged.
* **Hard pages: the hosted fetcher did not win anywhere.**
  * **5 of 6 were blocked** for Amazap too (`site_blocked`, all auto-credited), including **Nike, which curl fetches fine (HTTP 200)**.
  * The one success (TanStack docs) is *worse* than curl+trafilatura: the only code example is missing, and the headings are hoisted to the top with a stray `#`.
* **Economics are unchanged.**
  * vs raw HTML: pays on every page at $2+/MTok.
  * vs text-only: about break-even (single read at $4/MTok; 5 uncached turns at $2/MTok: 10/20 pages, +$0.30).
  * vs trafilatura: does not pay.

![chart](chart-v4.png)

## 20-page set: tokens and recall

| | Raw HTML | Text-only | trafilatura | Amazap v1 | v2 | v3 | **v4** |
|---|---:|---:|---:|---:|---:|---:|---:|
| Median tokens | 107,522 | 3,436 | 1,954 | 924 | 2,726 | 1,614 | **1,562** |
| Total tokens (20) | 2,826,066 | 109,186 | 69,097 | – | 117,379 | 58,464 | **57,646** |
| Median clean recall vs trafilatura | – | 98% | (reference) | 2% | 94% | 94% | **94%** |

For the Wikipedia pages, the body recall in brackets excludes trafilatura's reference-list lines, which article mode drops by design: #1 83%, #2 82%.

| # | Category | HTML | Text-only | trafilatura | Amazap v3 | **Amazap v4** | v4 vs HTML | v4 ÷ trafilatura | Recall clean (body*) | Text-only recall |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| 1 | wikipedia | 67,473 | 4,487 | 2,338 | 1,074 | **1,049** | -98.5% | 0.45× | 51% (83%) | 100% |
| 2 | wikipedia | 335,142 | 30,188 | 22,668 | 10,373 | **9,949** | -97.0% | 0.44× | 48% (82%) | 98% |
| 3 | docs-mdn | 48,158 | 3,188 | 2,178 | 2,401 | **2,400** | -95.0% | 1.10× | 94% | 94% |
| 4 | docs-mdn | 56,644 | 2,986 | 507 | 540 | **540** | -99.1% | 1.06× | 91% | 96% |
| 5 | docs-python | 32,574 | 7,441 | 5,876 | 6,643 | **6,640** | -79.6% | 1.13× | 96% | 99% |
| 6 | docs-python | 34,152 | 8,620 | 6,942 | 7,246 | **7,257** | -78.8% | 1.04× | 96% | 100% |
| 7 | docs-stripe | 597,196 | 4,597 | 3,314 | 2,386 | **2,384** | -99.6% | 0.72× | 83% | 88% |
| 8 | docs-stripe | 390,883 | 5,355 | 4,554 | 4,674 | **4,666** | -98.8% | 1.02× | 94% | 99% |
| 9 | github-readme | 123,793 | 1,555 | 398 | 554 | **477** | -99.6% | 1.20× | 81% | 95% |
| 10 | github-readme | 111,507 | 2,093 | 1,018 | 1,153 | **1,159** | -99.0% | 1.14× | 93% | 98% |
| 11 | news-article | 108,615 | 1,787 | 837 | 839 | **839** | -99.2% | 1.00× | 96% | 98% |
| 12 | news-hub | 183,713 | 2,067 | 462 | 551 | **532** | -99.7% | 1.15× | 98% | 96% |
| 13 | news-article | 63,196 | 1,716 | 1,009 | 1,039 | **1,033** | -98.4% | 1.02× | 99% | 97% |
| 14 | news-hub | 247,400 | 3,684 | 1,730 | 2,075 | **1,966** | -99.2% | 1.14× | 90% | 93% |
| 15 | blog | 22,517 | 10,082 | 8,496 | 9,501 | **9,439** | -58.1% | 1.11× | 99% | 100% |
| 16 | blog | 106,428 | 6,235 | 3,049 | 3,058 | **3,058** | -97.1% | 1.00× | 100% | 100% |
| 17 | product | 149,186 | 10,006 | 2,261 | 2,718 | **2,718** | -98.2% | 1.20× | 71% | 97% |
| 18 | gov | 9,762 | 648 | 177 | 264 | **182** | -98.1% | 1.03× | 100% | 100% |
| 19 | gov | 20,358 | 1,040 | 411 | 464 | **447** | -97.8% | 1.09× | 97% | 99% |
| 20 | dev-tool | 117,369 | 1,411 | 872 | 911 | **911** | -99.2% | 1.04× | 94% | 94% |

## 20-page set: headings, code, junk

| # | Page | h2/h3 in main (body only) | v4 kept (body) | v3 kept | trafilatura-md kept | Code: `<pre>` / v4 fences / traf fences | Citation tok | Ref-list tok | Boilerplate tok | **Junk share** | URL share | Title leaks | Empty `[](url)` links |
|---|---|---:|---:|---:|---:|---|---:|---:|---:|---:|---:|---:|---:|
| 1 | en.wikipedia.org/wiki/Lightning_Network | 8 (7) | 88% (100%) | 88% | 100% | 0 / 0 / 0 | 0 | 0 | 14 | **1.3%** | 0.0% | 0 | 0 |
| 2 | en.wikipedia.org/wiki/Large_language_model | 48 (44) | 92% (100%) | 92% | 98% | 0 / 0 / 0 | 0 | 0 | 279 | **2.8%** | 0.0% | 0 | 0 |
| 3 | developer.mozilla.org/en-US/docs/Web/JavaS | 15 (15) | 100% (100%) | 100% | 100% | 18 / 18 / 18 | 0 | 0 | 30 | **1.3%** | 0.0% | 0 | 0 |
| 4 | developer.mozilla.org/en-US/docs/Web/HTTP/ | 5 (5) | 100% (100%) | 100% | 100% | 3 / 3 / 3 | 0 | 0 | 0 | **0.0%** | 0.0% | 0 | 0 |
| 5 | docs.python.org/3/library/json.html | 11 (11) | 100% (100%) | 100% | 100% | 15 / 15 / 11 | 12 | 0 | 74 | **1.3%** | 0.0% | 0 | 0 |
| 6 | docs.python.org/3/tutorial/datastructures. | 13 (13) | 100% (100%) | 100% | 100% | 40 / 40 / 41 | 11 | 0 | 0 | **0.2%** | 0.0% | 0 | 0 |
| 7 | docs.stripe.com/api/charges/create | 6 (6) | 100% (100%) | 100% | 100% | 6 / 4 / 0 | 0 | 0 | 22 | **0.9%** | 0.0% | 0 | 0 |
| 8 | docs.stripe.com/webhooks | 27 (27) | 96% (96%) | 85% | 100% | 7 / 7 / 6 | 0 | 0 | 17 | **0.4%** | 0.0% | 0 | 0 |
| 9 | github.com/psf/requests | 5 (5) | 60% (60%) | 60% | 0% | 4 / 4 / 1 | 0 | 0 | 0 | **0.0%** | 0.0% | 0 | 0 |
| 10 | github.com/openai/tiktoken | 4 (4) | 100% (100%) | 25% | 0% | 6 / 6 / 6 | 0 | 0 | 0 | **0.0%** | 0.0% | 1 | 0 |
| 11 | www.bbc.com/news/articles/c3zxjdw5ywgpo | 1 (1) | 100% (100%) | 100% | 100% | 0 / 0 / 0 | 0 | 0 | 22 | **2.6%** | 0.0% | 0 | 0 |
| 12 | www.bbc.com/news/technology | 44 (44) | 100% (100%) | 100% | 100% | 0 / 0 / 0 | 0 | 0 | 2 | **0.4%** | 0.0% | 0 | 0 |
| 13 | www.npr.org/2026/09/23/nx-s1-5973306/ai-sl | 0 (0) | n/a (n/a) | n/a | n/a | 0 / 0 / 0 | 0 | 0 | 37 | **3.6%** | 0.0% | 0 | 0 |
| 14 | www.theverge.com/tech | 0 (0) | n/a (n/a) | n/a | n/a | 0 / 0 / 0 | 0 | 0 | 259 | **13.2%** | 0.0% | 0 | 0 |
| 15 | simonwillison.net/2024/Dec/31/llms-in-2024 | 2 (2) | 50% (50%) | 50% | 100% | 1 / 1 / 1 | 0 | 0 | 618 | **6.5%** | 0.0% | 0 | 0 |
| 16 | blog.cloudflare.com/the-road-to-quic/ | 9 (9) | 89% (89%) | 89% | 89% | 0 / 0 / 0 | 0 | 0 | 0 | **0.0%** | 0.0% | 0 | 0 |
| 17 | www.apple.com/macbook-air/ | 5 (5) | 80% (80%) | 80% | 0% | 0 / 0 / 0 | 0 | 0 | 565 | **20.8%** | 0.0% | 0 | 0 |
| 18 | www.usa.gov/passport | 4 (4) | 100% (100%) | 100% | 0% | 0 / 0 / 0 | 0 | 0 | 0 | **0.0%** | 0.0% | 0 | 0 |
| 19 | www.gov.uk/browse/driving | 18 (18) | 100% (100%) | 100% | 6% | 0 / 0 / 0 | 0 | 0 | 19 | **4.3%** | 0.0% | 0 | 0 |
| 20 | nodejs.org/en/about | 6 (6) | 100% (100%) | 100% | 100% | 2 / 2 / 2 | 0 | 0 | 15 | **1.6%** | 0.0% | 0 | 0 |

## Money (20-page set)

| Model ($/MTok) | Break-even tokens: 1 read / 5 turns / 5 cached | vs HTML 1 read | vs text 1 read | vs text 5 turns | vs text 5 cached | vs trafilatura 1 read | vs trafilatura 5 turns | vs trafilatura 5 cached |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| Claude Opus 5.5 ($4) | 2,689 / 538 / 1,855 | $+10.86 (20/20) | $-0.01 (4/20) | $+0.82 (18/20) | $+0.08 (6/20) | $-0.17 (1/20) | $+0.01 (3/20) | $-0.15 (1/20) |
| Claude Sonnet 5.5 ($2) | 5,379 / 1,076 / 3,710 | $+5.32 (20/20) | $-0.11 (2/20) | $+0.30 (10/20) | $-0.07 (2/20) | $-0.19 (1/20) | $-0.10 (2/20) | $-0.18 (1/20) |
| GPT-6.1 Sol ($2) | 5,379 / 1,076 / 3,710 | $+7.97 (20/20) | $-0.11 (2/20) | $+0.30 (10/20) | $-0.07 (2/20) | $-0.19 (1/20) | $-0.10 (2/20) | $-0.18 (1/20) |
| Gemini 3.1 Pro Preview ($2) | 5,379 / 1,076 / 3,842 | $+8.46 (20/20) | $-0.11 (2/20) | $+0.30 (10/20) | $-0.07 (2/20) | $-0.19 (1/20) | $-0.10 (2/20) | $-0.18 (1/20) |
| Claude Fable 5.1 (premium bracket) ($10) | 1,076 / 215 / 797 | $+27.47 (20/20) | $+0.30 (10/20) | $+2.36 (20/20) | $+0.48 (13/20) | $-0.10 (2/20) | $+0.36 (3/20) | $-0.06 (3/20) |
| Claude Haiku 5.5 (cheap bracket) ($0.1) | 107,577 / 21,515 / 65,198 | $+1.05 (11/20) | $-0.21 (0/20) | $-0.19 (0/20) | $-0.21 (0/20) | $-0.21 (0/20) | $-0.21 (0/20) | $-0.21 (0/20) |

Break-even uses the $0.010758 average fee. Each cell shows the net USD across the 20 pages, with the number of pages that save money in brackets. The median page saves **1,013 tokens** vs text-only (v3: 975) and is **55 tokens larger** than trafilatura. At $2/MTok one read needs 5,379 tokens of saving, 5 uncached turns need 1,076 per turn, and 5 cached turns (1.45×) need 3,710.

## Hard pages (separate; not in the medians above)

The brief: where a hosted fetcher might win. I probed candidates with curl first (same Chrome UA as the benchmark).
* Most "JS-heavy" sites are server-rendered and curl+trafilatura gets the content: TanStack, Tailwind, Vite, Zod, Figma, IKEA, Nike, CNN.
* True client-only apps (Excalidraw, the Netlify dashboard) have no article content to extract.
* So the hard set is mostly bot-protected pages, plus one SPA docs page and one JS product page.

| ID | Type | URL | curl status | curl+trafilatura gets | Amazap result | Order id | Charged / credited |
|---|---|---|---|---|---|---|---|
| H1 | spa-docs | https://tanstack.com/query/latest/docs/framework/react/overview | 200 | 396 words: "TanStack Query (formerly known as React Query) is often described as t…" | **delivered**, 589 words / 719 tok (recall vs trafilatura 77%) | `amazap_mcp_mv2a13vt_t8gjaf` | 13 / 0 |
| H2 | cloudflare-403 | https://www.npmjs.com/package/react | 403 | 6 words: "Enable JavaScript and cookies to continue…" | **site_blocked** ("The site blocked our fetcher.") | `amazap_mcp_mv2a18g9_k0kjt5` | 13 / 13 |
| H3 | curl-403 | https://stackoverflow.com/questions/11227809/why-is-processing-a-sorted-array-faster-than-processing-an-unsorted-array | 403 | 6 words: "Enable JavaScript and cookies to continue…" | **site_blocked** ("The site blocked our fetcher.") | `amazap_mcp_mv2a1e3p_b8761u` | 13 / 13 |
| H4 | js-product | https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111 | 200 | 160 words: "Nike Air Force 1 '07 Men's Shoes $115 Comfortable, durable and timeles…" | **site_blocked** ("The site blocked our fetcher.") | `amazap_mcp_mv2a1gqd_ks8gye` | 13 / 13 |
| H5 | news-js-datadome | https://www.reuters.com/business/ai-surveillance-startup-flock-safety-cut-several-hundred-jobs-amid-backlash-2026-10-09/ | 401 | 8 words: "Please enable JS and disable any ad blocker…" | **site_blocked** ("The site blocked our fetcher.") | `amazap_mcp_mv2a1jdy_z2nmh7` | 13 / 13 |
| H6 | curl-403-news-hub | https://apnews.com/hub/technology | 403 | 6 words: "Enable JavaScript and cookies to continue…" | **site_blocked** ("The site blocked our fetcher.") | `amazap_mcp_mv2a31hc_yn190k` | 13 / 13 |

**Verdict: no wins for the hosted fetcher on this set.**
* **4 pages block both:** npm (Cloudflare challenge), Stack Overflow, AP News and Reuters (DataDome). Curl gets a 403/401 challenge page ("Enable JavaScript and cookies to continue" / "Please enable JS and disable any ad blocker"); Amazap returns `site_blocked`. The buyer loses nothing, because all were auto-credited as **service credits** (not withdrawable sats) and the receipt says "Failed calls are not refunded in sats or cash".
* **Nike (H4):** curl gets **HTTP 200** and trafilatura extracts 160 words ("Nike Air Force 1 '07 Men's Shoes $115 Comfortable, durable and timeless…"). **Amazap is blocked.** Its fetcher is *more* blockable than plain curl here, probably because of its IP range or fingerprint.
* **TanStack Query docs (H1, React SPA with SSR):** both get the content.
  * Amazap: 719 tokens / 589 words; recall vs trafilatura 77%.
  * **Amazap is worse.** It drops the only code example: trafilatura has the full `QueryClientProvider` / `useQuery` snippet, and Amazap leaves just "Open in StackBlitz".
  * The three h2s are hoisted to the top with a trailing anchor `#`: `## You talked me into it, so what now?#`, `## Enough talk, show me some code already!#`, `## Motivation#`. They appear before the prose instead of in their sections.

## Verification of v3's "still broken" items

| v3 item | v4 status | Evidence (v4 markdown) |
|---|---|---|
| Link-title leaks (35 on Wikipedia LLM) | **Fixed** | 0 on all 20 pages. v3 had `("Chinchilla scaling "Chinchilla (language model)")")`; v4 has `("Chinchilla scaling")`. The single regex hit on #10 is a false positive (real quoted text: `"ing" (instead of e.g. "enc" and "oding")`). |
| Empty `[](url)` links (GitHub badges, Wikipedia files, share buttons) | **Fixed** | 0 empty links and **0 markdown links of any kind**, apart from the Python-docs footnote markers below. URL share is 0% on all 20 pages. GitHub requests dropped from 554 to 477 tokens. |
| Indented text → code blocks (usa.gov, gov.uk) | **Fixed** | usa.gov: "## Apply for a new adult passport / You need a passport to travel to most countries…", no indent. gov.uk: "### Driving licences / Apply for, renew or update your licence…", no indent. |
| usa.gov share block + stray `Top` | **Fixed** | usa.gov now ends at "…depends on if you are inside or outside the U.S." (182 tokens, from 264; trafilatura 177). No `Top` line on any page. |
| Stripe duplicate `-u` fence (#7) | **Fixed** | 4 fences: `bash` (5 lines), `response` (83), `bash` (2), `response` (77). The duplicate is gone, and curl fences are now tagged `bash`. |
| Stripe webhooks headings (#8, 23/27) | **Mostly fixed** | 26/27 (96%) of main-content h2/h3 are kept. |
| NPR duplicated photo credit | **Fixed** | A single line: "…trying to catch up. **Dan Kitwood/Getty Images**". |
| Wikipedia inline math | **Not fixed** | #2 still reads "…states that: where the variables are / and the statistical hyper-parameters are / - , meaning that it costs 6 FLOPs per parameter…". The formula and the variable-definition items are missing (trafilatura has the items "size of the artificial neural network itself…", "…size of its pretraining dataset…"). |
| Nav/caption trimming vs trafilatura | **Marginal** | Larger than trafilatura on 17/20 pages, now by a median of 55 tokens (v3: 88). Boilerplate tokens are nearly unchanged: Simon Willison 618 (651), Verge 259 (259), Apple 565 (565; mostly real product copy that trafilatura misses), BBC 22. Gains came from usa.gov (−82 tokens), Wikipedia LLM (−424) and GitHub (−77). |
| `quote_required` docs | **Partly fixed** | The error now includes a fresh quote to retry with: "Pass quote_id … A fresh quote is in quote. Retry buy with quote.quoteId. Nothing was charged." (`q_bf50172a`, used for H6). The `buy` description still says "quote is optional", and `quote_id` is still not in the schema's `required`. |
| Parallel `balanceAfterSats` | **Fixed** | Distinct and consistent (see above). |
| MCP inline cap (~24k chars) | **Unchanged** | 4/20 previews: #2 `amazap_mcp_mv29y4hi_h7m5o7`, #5 `amazap_mcp_mv29yidl_x39gdn`, #6 `amazap_mcp_mv29ynyu_trs8sd`, #15 `amazap_mcp_mv29zzd8_xz8gcg`. All signed `downloadUrl`s returned HTTP 200 with the full body. |

## Still broken (with order ids)

1. **Site blocks:**
   * Nike, which curl fetches: `amazap_mcp_mv2a1gqd_ks8gye`
   * npm: `amazap_mcp_mv2a18g9_k0kjt5`
   * Stack Overflow: `amazap_mcp_mv2a1e3p_b8761u`
   * Reuters: `amazap_mcp_mv2a1jdy_z2nmh7`
   * AP News: `amazap_mcp_mv2a31hc_yn190k`

   All were credited. The product currently gives no access advantage over curl on protected sites.
2. **TanStack Query docs:** code example dropped, and headings hoisted to the top with a trailing `#` (`amazap_mcp_mv2a13vt_t8gjaf`).
3. **Wikipedia math:** formulas and the variable-definition list are missing (`amazap_mcp_mv29y4hi_h7m5o7`).
4. **Python-docs footnote markers survive article mode:** `[\[1\]](#rfc-errata)` on #5 (`amazap_mcp_mv29yidl_x39gdn`) and `[\[1\]](#id2)` on #6 (`amazap_mcp_mv29ynyu_trs8sd`). These are small (≈12 tokens each).
5. **Credited receipts omit `balanceAfterSats`,** and the receipt vs `balance` session counters disagree (cosmetic).
6. **Inline cap ~24k chars:** 4/20 pages need a second fetch (ids above).
7. **Still slightly larger than trafilatura on 17/20 pages** (median +55 tokens). This is mostly kept headings, hub lists and some captions.

## Updated defensible blog claims

1. "Amazap Page to Markdown (article mode) returns **98% fewer tokens than raw HTML and half as many as a plain-text dump**, while keeping a median 94% of the main text, 100% of body headings and essentially every code block. Across 20 diverse pages it uses **17% fewer tokens in total than trafilatura**, though it is slightly larger (≈5%) on a typical page."
2. "The output is clean: zero link URLs, zero citation markers on Wikipedia, and language-tagged code fences. Receipts are exact: in our final run, 26 calls, 3 of them concurrent, reconciled to the sat. Blocked pages were auto-credited."
3. "At 13 sats (~$0.011) it pays for itself against raw HTML on every page with $2+/MTok models. Against a plain-text fetch it roughly breaks even: on one read at $4/MTok, and over ~5 uncached turns at $2/MTok. Against a local extractor it doesn't save money. You buy it for convenience and output quality."

**Do not claim** that it gets past bot protection or JS walls. On our 6 hard pages it was blocked on 5, including one that plain curl fetched.

## Where it does not pay off

* vs trafilatura: at most 3/20 pages in any scenario. Net is negative except at $4–10/MTok over 5 uncached turns (+$0.01 / +$0.36, driven by Wikipedia).
* vs text-only on a single read below ~$4/MTok, and with prompt caching at $2/MTok (2/20 pages).
* Haiku-class models: never vs text or trafilatura.
* Small pages.
* Bot-protected sites: blocked, so you get credits rather than content.

## Caveats

* n = 20 (+6 hard pages), one session, o200k_base.
* List prices.
* trafilatura is both the baseline and the recall reference. Wikipedia body recall excludes reference items.
* Junk is heuristic, and Apple's "boilerplate" is mostly real copy that trafilatura misses.
* Hard-page selection is limited by what was reachable today. Most "JS-heavy" sites are server-rendered, so the fetch problem today is bot protection rather than JS rendering.

## Files

* `report-v4.md`, `data-v4.csv` (20 pages), `data-v4-hard.csv` (6 hard pages), `ledger-v4.csv` (receipt-by-receipt balance check), `costs-v4.csv`, `summary-v4.json`, `chart-v4.png`
* `benchmark_v4.py`, `make_report_v4.py`, `report-v4.template.md`
* `urls_hard.txt`, `html_v4/`, `text_v4/`, `md_v4/`, `raw_buy_v4/`
