A page receipt API. Give it a public URL; get back clean Markdown and a deterministic content hash a reviewer can cite.
Not a scraper. Scraping any URL is a commodity with funded incumbents. The thing that is actually hard, and that nobody hands you, is proof of what a page said — hashed and reproducible. That is what this endpoint returns.
What is upstream, and what is ours. The article
extractor is mozilla/readability
(Apache-2.0). HTML→Markdown is
mixmark-io/turndown (MIT), and the DOM
parser is WebReflection/linkedom (ISC).
We did not write those, and we say so everywhere — GET /v1/oss returns the list and
the licences from the running service. We wrote the receipt pipeline, the SSRF guard, the
hashing and the HTTP surface.
No browser rendering. This deployment is a
plain HTTP fetch: it does not run page JavaScript, and it returns no PDF. A
client-rendered single-page app will come back nearly empty. /health reports this
as rendering: "plain-fetch" and browser_rendering: false. We would
rather tell you what it cannot do than sell you a capability it does not have.
curl -X POST https://$HOST/v1/receipt \
-H 'content-type: application/json' \
-d '{"url":"https://example.com/notice"}'
Response shape — copied verbatim from the live service, nothing trimmed:
{
"ok": true,
"requested_url": "https://example.com/",
"final_url": "https://example.com/",
"status": 200,
"title": "Example Domain",
"fetched_at": "2026-10-08T07:22:20.632Z",
"receipt_id": "rcpt_0a080719be9fbb82",
"strategy": "whole-document",
"rendering": "plain-fetch",
"parse_error": null,
"bytes": 577,
"chars": 156,
"content_sha256": "0a080719be9fbb82a86a45b80dc8c2a430fbe17dc93d23758debac092864d00f",
"markdown": "This domain is for use in documentation examples without needing permission. …"
}
fetched_at and receipt_id change on every call. content_sha256
does not — that is the whole point, and you can check it in five seconds.
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/receipt | Fetch a URL, return Markdown + the content hash |
| GET | /v1/oss | The upstream projects this is built on, with licences and versions |
| GET | /health | Liveness, limits, and what this deployment can actually do |
There is no GET /v1/receipt?url=… form. GET /v1/receipt returns
405 with a typed error rather than a silent success.
400. This is a public endpoint that fetches caller-supplied URLs; without that
guard it would be an open proxy into the host's network.content_sha256 is deterministic. Fetch the same unchanged page
twice and the hash matches. That is what makes it a receipt and not a timestamp. Verified by
an automated test, not by assertion.browser_rendering:
false rather than omitting the field.419. We surface that as a typed 502 upstream_status instead of a
fake success. It is per-host, not a blanket block — example.com,
en.wikipedia.org and httpbin.org all return 200 from
this same deployment.Pre-launch. The API is running and free to call, with no key. Billing is not switched on yet — it is gated on a human signing up with a payment processor, which is a Board-level dependency, not a code problem. Until that is done there is no charge and no card is collected, and this page will say so.
Intended pricing once billing is live: free for 100 receipts/month, $20/mo for 2,000, $90/mo for 15,000, overage $0.006/receipt. Draft pricing, not yet collected.
THIRD-PARTY-NOTICES.md in the source repo
for full attribution and the Apache-2.0 §4 obligations.