Maya

AWS CloudFront

Stream CloudFront access logs to Maya through an Amazon Data Firehose HTTP endpoint — no code, no Lambda. Maya keeps only AI-agent hits.

Audience: Engineering, AI agentsUpdated 2026-10-06
Claude CodeCodexGitHub Copilot
Set this up with your coding agent

Copy the prompt and paste it into Cursor, Claude Code, Codex, Copilot… — it does the install for you.

What it is

A no-code pipeline that streams your CloudFront access logs to Maya and turns them into AI-crawler visibility — which bots (GPTBot, ClaudeBot, PerplexityBot, …) crawl which pages, how often, and where they hit errors:

Text
CloudFront distribution
  → Standard logging v2 (output format: JSON)
  → Amazon Data Firehose (HTTP endpoint destination)
  → Maya ingest endpoint (authenticated with your access key)

Everything runs on AWS-managed services — there is no Lambda to write, no log shipper to run, and Maya never receives AWS credentials. Maya hands you two values (an endpoint URL and an access key); you paste them into a Firehose stream once.

Maya parses each delivered log line, keeps only requests from known AI agents, and drops everything else. Human traffic is never stored.

Prerequisites

  • A CloudFront distribution serving the domain of your Maya project.
  • An AWS account with permission to create a Firehose stream and edit the distribution.
  • The CloudFront integration enabled in Maya: Crawler Visibility → Settings → AWS CloudFront → Enable Monitoring. This generates your HTTP endpoint URL and access key.

Step 1 — Create the Firehose stream

Region requirement. CloudFront only delivers standard logs to a Firehose stream in US East (N. Virginia) us-east-1. Streams in any other region will not appear as a valid destination.

  1. Open Amazon Data Firehose in us-east-1 → Create Firehose stream.
  2. Source: Direct PUT. Destination: HTTP endpoint.
  3. HTTP endpoint URL: paste the URL from Maya settings (it already contains your project id).
  4. Access key: paste the access key from Maya settings. Firehose sends it as the X-Amz-Firehose-Access-Key header on every delivery.
  5. Content encoding: GZIP (optional, recommended — Maya accepts both).
  6. Buffer hints: set buffer size to 1 MiB and buffer interval to 60 seconds. Keep the buffer small — Maya's ingest accepts request bodies up to a few MB, and a small buffer also means fresher data.
  7. Leave retry duration at the default and (optionally) configure the S3 backup bucket for failed deliveries.

Step 2 — Enable standard logging v2 on the distribution

  1. CloudFront → your distribution → Logging tab → Add → Firehose.
  2. Destination: the stream you just created.
  3. Output format: JSON. This matters — Maya parses JSON log lines; plain/w3c output is rejected with a hint in the Maya settings card. The output format is fixed when the delivery destination is created; to change it you must delete and re-create the delivery.
  4. Field selection: include at least date, cs(Host), cs-uri-stem, sc-status, sc-bytes, cs(User-Agent). Extra fields are fine — Maya ignores what it doesn't need.

Step 3 — Verify

  • CloudFront delivers logs within minutes of traffic. After the first delivery, the Maya settings card shows a Last delivery timestamp.
  • AI-crawler hits appear in the Crawler Visibility dashboard tagged with source cloudfront.
  • No traffic yet? Hit a page of your site with a bot user agent, e.g.:
Bash
curl -A "Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)" https://your-domain.com/

How delivery works

  • Authentication. Every delivery carries your access key in the X-Amz-Firehose-Access-Key header. Regenerating the key in Maya invalidates the old one immediately — update the Firehose destination afterwards.
  • At-least-once delivery. Firehose retries failed batches with the same request id; Maya deduplicates on that id, so retries never inflate your numbers.
  • Filtering. Static assets (.js, .css, images, fonts) and internal paths are not counted as page crawls. Requests whose user agent doesn't match a known AI agent are discarded.

Troubleshooting

SymptomCause / fix
Settings card shows "Received log lines that are not JSON"The delivery destination's output format is plain/w3c. Delete the delivery and re-create it with output format JSON.
Firehose shows delivery failures with HTTP 401Access key mismatch — re-copy the key from Maya settings into the Firehose destination (it changes if you clicked Regenerate Key).
No Last delivery timestampCheck the stream exists in us-east-1, the distribution's Logging tab shows the Firehose delivery as Enabled, and the site actually received traffic since enabling (logging changes can take up to ~12 h to take effect).
Deliveries fail with HTTP 413The delivery batch is too large. Lower the Firehose buffer size to 1 MiB and enable GZIP content encoding.

Pricing note

CloudFront doesn't charge for enabling standard logs, but the delivery is billed as CloudWatch Vended Logs plus standard Firehose ingestion. For AI-crawler volumes this is typically cents per month; see the Vended Logs section of CloudWatch pricing.

Alternatives

  • Already filtering at the edge? The universal streaming endpoint accepts any NDJSON access-log stream (Cloudflare Logpush, Vector, Fluent Bit).
  • For the highest-privacy path where Maya never receives a non-bot line, see batch log export.