## [](#what-it-is)What it is

A single, platform-agnostic endpoint that accepts your access logs as they happen:

Text

```
POST https://logs.withmaya.ai/v1/logs
```

Whatever platform you run — Cloudflare, Vercel, a CDN, or a plain Nginx box — you point an access-log stream you already produce at this endpoint. Maya normalizes each line, classifies the AI agent behind it, and surfaces it in your Signals dashboard. There is no per-platform script to write; the endpoint does the parsing.

This is the **fastest** way to see AI-agent traffic. For the **highest-privacy** path, where you filter at the source and Maya never receives a non-bot line, see the [batch log export](./overview.md) — the two modes run side by side and produce the same dashboards.

## [](#how-it-differs-from-batch-export)How it differs from batch export

Real-time streaming (this page)

[Batch export](./overview.md)

Delivery

HTTP POST as requests happen

Daily/weekly upload or mTLS pull

Who filters

Maya's edge worker (`classifyAgent`) keeps agents; humans are dropped

The brand filters at the source before sending

Latency

Seconds

Hours to a day

Setup

Point a log stream at one URL

Deploy a reference filter next to log rotation

Privacy posture

Maya's edge sees each streamed line in order to classify it

Maya never receives a non-allowlisted line

> **Choosing.** If your platform can pre-filter (e.g. a Cloudflare Logpush job scoped to bot traffic), streaming stays privacy-preserving _and_ real-time. If you need the hard guarantee that Maya never sees a human request, use batch export. Regulated brands (bank IT, strict DPAs) should default to batch; everyone else gets the fastest result from streaming.

## [](#request-contract)Request contract

HTTP

```
POST https://logs.withmaya.ai/v1/logs
Content-Type: application/x-ndjson
x-maya-signal-project: prj_...
x-maya-signal-key: msig_live_...
x-maya-signal-source: src_...        (required — see note)
x-maya-signal-zone: example.com      (optional; otherwise derived from each line's host)
```

*   **Body:** newline-delimited JSON (the common log format), a JSON array, or `{ "records": [...] }`. Up to **5,000 records** per request.
*   **Filtering:** by default only AI-agent visits are forwarded — the reason this pipeline exists. Append `?all=1` to keep every request for full server-side analytics.
*   **Response:** `202` with `{ received, agent_events, forwarded, rejected }`. A nonzero `rejected` with `forwarded: 0` means the key or scope was refused downstream — check the scope note below.
*   A line without a request method or User-Agent is skipped.

> **Scope note.** The streaming endpoint does not validate the write key itself — it forwards each normalized event to Maya's ingest, which enforces scope. A production `msig_live_` key is bound to one source (and, if configured, a hostname allowlist). So: **send `x-maya-signal-source`** — without it the event has no source and is rejected `403`. And make sure `x-maya-signal-zone` (or each line's host) matches an allowed hostname. Your project and source IDs are permanent — set them once; they do not change between sessions.

## [](#field-mapping)Field mapping

You do not have to reshape your logs into a Maya schema. Each line is normalized into a neutral field set, and Maya's own names, Cloudflare Logpush names, and common access-log aliases are all accepted:

Normalized

Accepted keys

method

`method`, `ClientRequestMethod`

host

`host`, `ClientRequestHost`, `EdgeRequestHost`

url / path + query

`url`, `ClientRequestURI`, `uri`, `path`

status

`status`, `EdgeResponseStatus`, `response_status`

user\_agent

`user_agent`, `userAgent`, `ClientRequestUserAgent`

referer

`referer`, `referrer`, `ClientRequestReferer`

response bytes

`response_bytes`, `EdgeResponseBytes`

cache status

`cache_status`, `CacheCacheStatus`

country

`country`, `ClientCountry`

asn

`asn`, `ClientASN`

bot score

`bot_score`, `BotScore`

verified bot

`verified_bot`, `VerifiedBot`, `VerifiedBotCategory`

timestamp

`timestamp`, `EdgeStartTimestamp`, `time`, `occurred_at`

Only whitelisted UTM parameters (`utm_*`, `gclid`, `fbclid`, `msclkid`) are read from the query string; the rest of the query, cookies, and authorization headers are discarded.

## [](#set-up-cloudflare-logpush)Set up Cloudflare Logpush

Cloudflare is the cleanest feeder: Logpush can filter to bot traffic at the edge and deliver over HTTP with custom headers.

1.  Cloudflare dashboard → your zone → **Analytics & Logs → Logpush → Create a Logpush job**, dataset **HTTP requests**.
    
2.  **Destination: HTTP.** Custom headers are passed as `header_*` query parameters on the destination URL:
    
    Text
    
    ```
    https://logs.withmaya.ai/v1/logs?header_x-maya-signal-project=prj_...&header_x-maya-signal-source=src_...&header_x-maya-signal-key=msig_live_...
    ```
    
3.  **(Recommended) Filter at the edge** so only bots leave Cloudflare — e.g. `BotScore lt 30` or a set `VerifiedBotCategory`. This keeps streaming privacy-preserving.
    
4.  **Select fields:** `ClientRequestMethod`, `ClientRequestHost`, `ClientRequestURI`, `EdgeResponseStatus`, `ClientRequestUserAgent`, `VerifiedBotCategory`, `BotScore`, `EdgeStartTimestamp`.
    
5.  Save the job. The Signals dashboard flips to **Connected** as the first agent hit arrives.
    

> The write key is shown only once, at issue or rotation. Copy it straight into the Logpush destination then. Your `prj_` and `src_` IDs never change, so the job never needs re-pointing.

Other platforms (Vercel Log Drains, AWS CloudFront, Fastly, Nginx via a shipper) post to the same endpoint with the same headers; they differ only in how the stream is configured.

## [](#example-transform)Example transform

Incoming Cloudflare Logpush line:

JSON

```
{
  "ClientRequestMethod": "GET",
  "ClientRequestHost": "shop.com",
  "ClientRequestURI": "/blog?utm_source=chatgpt&session=secret",
  "EdgeResponseStatus": 200,
  "ClientRequestUserAgent": "Mozilla/5.0 (compatible; PerplexityBot/1.0)",
  "VerifiedBotCategory": "AI Crawler"
}
```

Stored event:

JSON

```
{
  "source": "log-drain",
  "request": { "method": "GET", "path": "/blog", "attribution": { "utm_source": "chatgpt" } },
  "agent": { "name": "PerplexityBot", "request_type": "crawler", "verified": true }
}
```

The raw query (`session=secret`) is discarded; only whitelisted UTM is kept.

## [](#verification)Verification

*   **`202` with `forwarded > 0`:** events reached Maya. The dashboard updates within seconds.
*   **`202` with `rejected > 0, forwarded: 0`:** the key/scope was refused — confirm `x-maya-signal-source` is set and the host matches the source's allowlist.
*   **`agent_events: 0`:** the batch held no AI agents. Send `?all=1` to confirm the pipe works with human traffic, or check that your field mapping carries a User-Agent.
*   **`401`:** missing `x-maya-signal-project` or `x-maya-signal-key`.

## [](#privacy-contract)Privacy contract

*   Only whitelisted UTM parameters are read from any query string; the rest of the query, fragments, cookies, authorization headers, and raw IPs are never stored.
*   Server sources carry request/response metadata only — never bodies.
*   In streaming mode Maya's edge worker sees each streamed line to classify it. To keep Maya from ever seeing a non-bot line, filter at the source (Logpush edge filter) or use [batch export](./overview.md).

[PreviousCustom BFF Endpoint](/docs/integrations/log-export/bff-endpoint)[Next Overview](/docs/integrations/markdown-rendering/overview)

![](/assets/landing-v2/maya-outline.svg)

THE NEXT ANSWER COULD BE YOURS.

## Get your brand  
mentioned in AI Search.

[Let’s make it happen](https://withmaya.ai/demo)