AWS CloudFront
Stream CloudFront access logs to Maya through an Amazon Data Firehose HTTP endpoint — no code, no Lambda. Maya keeps only AI-agent hits.
Copy the prompt and paste it into Cursor, Claude Code, Codex, Copilot… — it does the install for you.
What it is
A no-code pipeline that streams your CloudFront access logs to Maya and turns them into AI-crawler visibility — which bots (GPTBot, ClaudeBot, PerplexityBot, …) crawl which pages, how often, and where they hit errors:
CloudFront distribution
→ Standard logging v2 (output format: JSON)
→ Amazon Data Firehose (HTTP endpoint destination)
→ Maya ingest endpoint (authenticated with your access key)Everything runs on AWS-managed services — there is no Lambda to write, no log shipper to run, and Maya never receives AWS credentials. Maya hands you two values (an endpoint URL and an access key); you paste them into a Firehose stream once.
Maya parses each delivered log line, keeps only requests from known AI agents, and drops everything else. Human traffic is never stored.
Prerequisites
- A CloudFront distribution serving the domain of your Maya project.
- An AWS account with permission to create a Firehose stream and edit the distribution.
- The CloudFront integration enabled in Maya: Crawler Visibility → Settings → AWS CloudFront → Enable Monitoring. This generates your HTTP endpoint URL and access key.
Step 1 — Create the Firehose stream
Region requirement. CloudFront only delivers standard logs to a Firehose stream in US East (N. Virginia)
us-east-1. Streams in any other region will not appear as a valid destination.
- Open Amazon Data Firehose in
us-east-1→ Create Firehose stream. - Source: Direct PUT. Destination: HTTP endpoint.
- HTTP endpoint URL: paste the URL from Maya settings (it already contains your project id).
- Access key: paste the access key from Maya settings. Firehose sends it as the
X-Amz-Firehose-Access-Keyheader on every delivery. - Content encoding: GZIP (optional, recommended — Maya accepts both).
- Buffer hints: set buffer size to 1 MiB and buffer interval to 60 seconds. Keep the buffer small — Maya's ingest accepts request bodies up to a few MB, and a small buffer also means fresher data.
- Leave retry duration at the default and (optionally) configure the S3 backup bucket for failed deliveries.
Step 2 — Enable standard logging v2 on the distribution
- CloudFront → your distribution → Logging tab → Add → Firehose.
- Destination: the stream you just created.
- Output format: JSON. This matters — Maya parses JSON log lines; plain/w3c output is rejected with a hint in the Maya settings card. The output format is fixed when the delivery destination is created; to change it you must delete and re-create the delivery.
- Field selection: include at least
date,cs(Host),cs-uri-stem,sc-status,sc-bytes,cs(User-Agent). Extra fields are fine — Maya ignores what it doesn't need.
Step 3 — Verify
- CloudFront delivers logs within minutes of traffic. After the first delivery, the Maya settings card shows a Last delivery timestamp.
- AI-crawler hits appear in the Crawler Visibility dashboard tagged with source
cloudfront. - No traffic yet? Hit a page of your site with a bot user agent, e.g.:
curl -A "Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)" https://your-domain.com/How delivery works
- Authentication. Every delivery carries your access key in the
X-Amz-Firehose-Access-Keyheader. Regenerating the key in Maya invalidates the old one immediately — update the Firehose destination afterwards. - At-least-once delivery. Firehose retries failed batches with the same request id; Maya deduplicates on that id, so retries never inflate your numbers.
- Filtering. Static assets (
.js,.css, images, fonts) and internal paths are not counted as page crawls. Requests whose user agent doesn't match a known AI agent are discarded.
Troubleshooting
| Symptom | Cause / fix |
|---|---|
| Settings card shows "Received log lines that are not JSON" | The delivery destination's output format is plain/w3c. Delete the delivery and re-create it with output format JSON. |
| Firehose shows delivery failures with HTTP 401 | Access key mismatch — re-copy the key from Maya settings into the Firehose destination (it changes if you clicked Regenerate Key). |
| No Last delivery timestamp | Check the stream exists in us-east-1, the distribution's Logging tab shows the Firehose delivery as Enabled, and the site actually received traffic since enabling (logging changes can take up to ~12 h to take effect). |
| Deliveries fail with HTTP 413 | The delivery batch is too large. Lower the Firehose buffer size to 1 MiB and enable GZIP content encoding. |
Pricing note
CloudFront doesn't charge for enabling standard logs, but the delivery is billed as CloudWatch Vended Logs plus standard Firehose ingestion. For AI-crawler volumes this is typically cents per month; see the Vended Logs section of CloudWatch pricing.
Alternatives
- Already filtering at the edge? The universal streaming endpoint accepts any NDJSON access-log stream (Cloudflare Logpush, Vector, Fluent Bit).
- For the highest-privacy path where Maya never receives a non-bot line, see batch log export.