Maya

Real-time Log Streaming

Point any access-log stream at a single Maya endpoint. Maya normalizes each line, keeps only AI-agent traffic, and shows it in your dashboard — no per-platform code.

Audience: Engineering, AI agentsUpdated 2026-09-21
Claude CodeCodexGitHub Copilot
Set this up with your coding agent

Copy the prompt and paste it into Cursor, Claude Code, Codex, Copilot… — it does the install for you.

What it is

A single, platform-agnostic endpoint that accepts your access logs as they happen:

Text
POST https://logs.withmaya.ai/v1/logs

Whatever platform you run — Cloudflare, Vercel, a CDN, or a plain Nginx box — you point an access-log stream you already produce at this endpoint. Maya normalizes each line, classifies the AI agent behind it, and surfaces it in your Signals dashboard. There is no per-platform script to write; the endpoint does the parsing.

This is the fastest way to see AI-agent traffic. For the highest-privacy path, where you filter at the source and Maya never receives a non-bot line, see the batch log export — the two modes run side by side and produce the same dashboards.

How it differs from batch export

Real-time streaming (this page)Batch export
DeliveryHTTP POST as requests happenDaily/weekly upload or mTLS pull
Who filtersMaya's edge worker (classifyAgent) keeps agents; humans are droppedThe brand filters at the source before sending
LatencySecondsHours to a day
SetupPoint a log stream at one URLDeploy a reference filter next to log rotation
Privacy postureMaya's edge sees each streamed line in order to classify itMaya never receives a non-allowlisted line

Choosing. If your platform can pre-filter (e.g. a Cloudflare Logpush job scoped to bot traffic), streaming stays privacy-preserving and real-time. If you need the hard guarantee that Maya never sees a human request, use batch export. Regulated brands (bank IT, strict DPAs) should default to batch; everyone else gets the fastest result from streaming.

Request contract

HTTP
POST https://logs.withmaya.ai/v1/logs
Content-Type: application/x-ndjson
x-maya-signal-project: prj_...
x-maya-signal-key: msig_live_...
x-maya-signal-source: src_...        (required — see note)
x-maya-signal-zone: example.com      (optional; otherwise derived from each line's host)
  • Body: newline-delimited JSON (the common log format), a JSON array, or { "records": [...] }. Up to 5,000 records per request.
  • Filtering: by default only AI-agent visits are forwarded — the reason this pipeline exists. Append ?all=1 to keep every request for full server-side analytics.
  • Response: 202 with { received, agent_events, forwarded, rejected }. A nonzero rejected with forwarded: 0 means the key or scope was refused downstream — check the scope note below.
  • A line without a request method or User-Agent is skipped.

Scope note. The streaming endpoint does not validate the write key itself — it forwards each normalized event to Maya's ingest, which enforces scope. A production msig_live_ key is bound to one source (and, if configured, a hostname allowlist). So: send x-maya-signal-source — without it the event has no source and is rejected 403. And make sure x-maya-signal-zone (or each line's host) matches an allowed hostname. Your project and source IDs are permanent — set them once; they do not change between sessions.

Field mapping

You do not have to reshape your logs into a Maya schema. Each line is normalized into a neutral field set, and Maya's own names, Cloudflare Logpush names, and common access-log aliases are all accepted:

NormalizedAccepted keys
methodmethod, ClientRequestMethod
hosthost, ClientRequestHost, EdgeRequestHost
url / path + queryurl, ClientRequestURI, uri, path
statusstatus, EdgeResponseStatus, response_status
user_agentuser_agent, userAgent, ClientRequestUserAgent
refererreferer, referrer, ClientRequestReferer
response bytesresponse_bytes, EdgeResponseBytes
cache statuscache_status, CacheCacheStatus
countrycountry, ClientCountry
asnasn, ClientASN
bot scorebot_score, BotScore
verified botverified_bot, VerifiedBot, VerifiedBotCategory
timestamptimestamp, EdgeStartTimestamp, time, occurred_at

Only whitelisted UTM parameters (utm_*, gclid, fbclid, msclkid) are read from the query string; the rest of the query, cookies, and authorization headers are discarded.

Set up Cloudflare Logpush

Cloudflare is the cleanest feeder: Logpush can filter to bot traffic at the edge and deliver over HTTP with custom headers.

  1. Cloudflare dashboard → your zone → Analytics & Logs → Logpush → Create a Logpush job, dataset HTTP requests.

  2. Destination: HTTP. Custom headers are passed as header_* query parameters on the destination URL:

    Text
    https://logs.withmaya.ai/v1/logs?header_x-maya-signal-project=prj_...&header_x-maya-signal-source=src_...&header_x-maya-signal-key=msig_live_...
  3. (Recommended) Filter at the edge so only bots leave Cloudflare — e.g. BotScore lt 30 or a set VerifiedBotCategory. This keeps streaming privacy-preserving.

  4. Select fields: ClientRequestMethod, ClientRequestHost, ClientRequestURI, EdgeResponseStatus, ClientRequestUserAgent, VerifiedBotCategory, BotScore, EdgeStartTimestamp.

  5. Save the job. The Signals dashboard flips to Connected as the first agent hit arrives.

The write key is shown only once, at issue or rotation. Copy it straight into the Logpush destination then. Your prj_ and src_ IDs never change, so the job never needs re-pointing.

Other platforms (Vercel Log Drains, AWS CloudFront, Fastly, Nginx via a shipper) post to the same endpoint with the same headers; they differ only in how the stream is configured.

Example transform

Incoming Cloudflare Logpush line:

JSON
{
  "ClientRequestMethod": "GET",
  "ClientRequestHost": "shop.com",
  "ClientRequestURI": "/blog?utm_source=chatgpt&session=secret",
  "EdgeResponseStatus": 200,
  "ClientRequestUserAgent": "Mozilla/5.0 (compatible; PerplexityBot/1.0)",
  "VerifiedBotCategory": "AI Crawler"
}

Stored event:

JSON
{
  "source": "log-drain",
  "request": { "method": "GET", "path": "/blog", "attribution": { "utm_source": "chatgpt" } },
  "agent": { "name": "PerplexityBot", "request_type": "crawler", "verified": true }
}

The raw query (session=secret) is discarded; only whitelisted UTM is kept.

Verification

  • 202 with forwarded > 0: events reached Maya. The dashboard updates within seconds.
  • 202 with rejected > 0, forwarded: 0: the key/scope was refused — confirm x-maya-signal-source is set and the host matches the source's allowlist.
  • agent_events: 0: the batch held no AI agents. Send ?all=1 to confirm the pipe works with human traffic, or check that your field mapping carries a User-Agent.
  • 401: missing x-maya-signal-project or x-maya-signal-key.

Privacy contract

  • Only whitelisted UTM parameters are read from any query string; the rest of the query, fragments, cookies, authorization headers, and raw IPs are never stored.
  • Server sources carry request/response metadata only — never bodies.
  • In streaming mode Maya's edge worker sees each streamed line to classify it. To keep Maya from ever seeing a non-bot line, filter at the source (Logpush edge filter) or use batch export.