Skip to main content

Http Parsing

The following services are managing HTTP communications parsing.
Below is their default configuration.

webWrite​

webWrite is the Helm values key for the web parser (web-parser, renamed from web-write / web-streams-write). Since the HTTP/2 work, it parses HTTP/1.x, HTTP/2 and gRPC together - there is no separate gRPC parser service. Its configuration reflects that: one parsingJob reading a single merged queue, and a protocols block carrying one store/cache set per protocol it may output.

"@id": web-parser
"@type": ServerConfiguration
version: '0.1'
protocols:
http:
comStore:
node: http://elasticsearch:9200
indexPurge: spider-search-httpcom-upload
indexGet: spider-search-httpcom
getTimeout: PT10S
purgeTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
contentStore:
node: http://elasticsearch:9200
indexGet: spider-search-httpcomcontent
indexPurge: spider-active-httpcomcontent-upload-default
getTimeout: PT15S
purgeTimeout: PT15S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT15S
circuitThreshold: 0.8
comsCache:
server: redis
port: 6379
db: 2
timeOut: PT10S
ttl: PT25S
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
comsContentCache:
server: redis
port: 6379
db: 12
timeOut: PT10S
ttl: PT25S
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
grpc: # present only when the grpcParsing feature flip is on
comStore:
node: http://elasticsearch:9200
indexPurge: spider-search-grpccom-upload
indexGet: spider-search-grpccom
getTimeout: PT10S
purgeTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
contentStore:
node: http://elasticsearch:9200
indexPurge: spider-search-grpccomcontent-upload
indexGet: spider-search-grpccomcontent
getTimeout: PT10S
purgeTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
comsCache: # a distinct Redis instance (redis-grpc), not the HTTP cache's db
server: redis-grpc
port: 6379
db: 2
timeOut: PT10S
ttl: PT25S
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
comsContentCache:
server: redis-grpc
port: 6379
db: 12
timeOut: PT10S
ttl: PT25S
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
parsingLog:
store:
node: http://elasticsearch:9200
indexPurge: spider-search-httppers-upload
indexGet: spider-search-httppers
getTimeout: PT10S
purgeTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
cache:
server: redis
port: 6379
db: 3
timeOut: PT10S
ttl: PT2M
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
parsingJob:
pollingDelay: PT1S
jobRunners: 1
uri: http://tcp-update/v1/parsing-jobs/web
label: "/tcp-update/parsing-jobs/web"
timeout: PT10S
parsingDelay: PT10S
queues: ["web", "http", "grpc"] # "grpc" dropped from the list when grpcParsing is off
purgeJob:
sizeLimitKB: 100
parsingStatusSynchro:
delay: PT5S
queue: statsSynchro
maxBucketSize: 100
tcpSessions:
uri: http://tcp-update/v1/tcp-sessions
label: "/tcp-update/tcp-sessions"
timeout: PT10S
sizeLimitKB: 100
compressed: true
packetsByIndex:
uri: http://pack-read/v1/packets/of/tcpsession
label: "/pack-read/packets/of/tcpsession"
timeout: PT10S
whisps:
get:
uri: http://whisp/v1/whisperers/{id}
label: "/whisp/whisperers"
timeout: PT10S
config:
uri: http://whisp/v1/whisperers/{id}/config?view=full
label: "/whisp/whisperers/config"
timeout: PT10S

There is one parsingLog (parsing state / tracking resource), shared across protocols - a TCP session being parsed for web traffic has a single parsing log regardless of whether it turns out to carry HTTP or gRPC.

Two related settings live in the shared common configuration, not in webWrite itself:

  • parserProtocols - the list of protocols with a dedicated parsing queue (drives queue naming, e.g. the web / http / grpc queues above).
  • outputProtocols - the list of protocols the web parser may emit a communication for (http, grpc, websocket, sse). Both sse and websocket have an engine and appear in data - see SSE and WebSocket message parsing settings.

webRead​

"@id": web-streams-read
"@type": ServerConfiguration
version: '0.1'
httpComStore:
node: http://elasticsearch:9200
indexGet: spider-search-httpcom
getTimeout: PT15S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT15S
circuitThreshold: 0.8
httpComContentStore:
node: http://elasticsearch:9200
indexGet: spider-search-httpcomcontent
getTimeout: PT15S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT15S
circuitThreshold: 0.8
httpPersStore:
node: http://elasticsearch:9200
indexGet: spider-search-httppers
indexParsingStatus: spider-search-parsing-status-httppers
getTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
packets:
uri: http://pack-read/v1/packets/payload/tcp/?force=true
label: "/packets/payload/tcp"
timeout: PT10S
searchRequest:
sizeLimitKB: 100
whisps:
uri: http://whisp/v1/whisperers/{id}/config?view=server
label: "/whisp/whisperers/config"
timeout: PT10S

webHttpComPoller​

"@id": web-httpcom-poller
"@type": ServerConfiguration
version: '0.1'
logField: httpCom
itemStore:
node: http://elasticsearch:9200
useDataStoragePolicies: true
useDistinctPolicyForUpload: true
streamedDataStoragePolicies:
- name: default
indexSave: spider-active-httpcom-streaming-default
uploadedDataStoragePolicies:
- name: default
indexSave: spider-active-httpcom-upload-default
saveTimeout: PT5S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
itemCache:
server: redis
port: 6379
db: 2
timeOut: PT2S
ttl: PT45S
circuitDuration: PT15S
circuitThreshold: 1
compressed: false
polling:
queue: httpComToSynchronize
queueType: SORTED_SET
scoreAttribute: _update
size: 175
jobRunners: 1
delay: PT2S
automaticESIds: false
updateCache:
keepInCache: false
removeFromCacheCondition: {}
noTTLCondition: {}
whisps:
config:
uri: http://whisp/v1/whisperers/{id}/config?view=full
label: "/whisp/whisperers/config"
timeout: PT10S

webHttpComContentPoller​

"@id": web-httpcom-content-poller
"@type": ServerConfiguration
version: '0.1'
logField: httpCom
itemStore:
node: http://elasticsearch:9200
useDataStoragePolicies: true
useDistinctPolicyForUpload: true
streamedDataStoragePolicies:
- name: default
indexSave: spider-active-httpcomcontent-streaming-default
uploadedDataStoragePolicies:
- name: default
indexSave: spider-active-httpcomcontent-upload-default
saveTimeout: PT5S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
itemCache:
server: redis
port: 6379
db: 12
timeOut: PT2S
ttl: PT45S
circuitDuration: PT15S
circuitThreshold: 1
compressed: false
polling:
queue: httpComContentToSynchronize
queueType: SORTED_SET
scoreAttribute: _update
size: 175
jobRunners: 1
delay: PT2S
automaticESIds: false
updateCache:
keepInCache: false
removeFromCacheCondition: {}
noTTLCondition: {}
whisps:
config:
uri: http://whisp/v1/whisperers/{id}/config?view=full
label: "/whisp/whisperers/config"
timeout: PT10S

webHttpPersPoller​

"@id": web-httppers-poller
"@type": ServerConfiguration
version: '0.1'
logField: httpPers
itemStore:
node: http://elasticsearch:9200
useDataStoragePolicies: true
useDistinctPolicyForUpload: true
streamedDataStoragePolicies:
- name: default
indexSave: spider-active-httppers-streaming-default
uploadedDataStoragePolicies:
- name: default
indexSave: spider-active-httppers-upload-default
saveTimeout: PT5S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
itemCache:
server: redis
port: 6379
db: 3
timeOut: PT2S
ttl: PT45S
longTtl: PT3M30S
circuitDuration: PT15S
circuitThreshold: 1
compressed: false
polling:
queue: httpPersToSynchronize
queueType: SORTED_SET_HMAP
scoreAttribute: first
size: 175
jobRunners: 1
delay: PT5S
automaticESIds: false
updateCache:
keepInCache: true
useLongTtlAsSafety: true
removeFromCacheCondition:
property: status
values:
- COMPLETED
- ERROR
noTTLCondition: {}
saveInESConditions:
- property: state
values:
- CLOSED
- CLOSE_WAIT
- ESTABLISHED
- LAST_ACK
whisps:
config:
uri: http://whisp/v1/whisperers/{id}/config?view=full
label: "/whisp/whisperers/config"
timeout: PT10S

GraphQL​

GraphQL is not a separate parser - it is an enrichment the web parser applies to the HTTP/1.x and HTTP/2 communications it already produces on the paths configured in server.graphql.paths (default ["/graphql"]). Off the configured paths, nothing is detected unless heuristic is turned on, which widens detection at the risk of mislabelling a REST body that happens to carry a query string key - it is off by default for that reason.

Detected shapes:

  • POST JSON - content-type: application/json, body an object with a string query, or an object carrying extensions.persistedQuery.sha256Hash (a persisted / APQ hash-only request), or a non-empty array of such objects (a batch).
  • POST raw - content-type: application/graphql, body is the document itself.
  • GET - the query string carries query= (and optionally operationName=, variables=, extensions=).

The body used is the decoded one (gzip/deflate/br undone), capped at server.graphql.maxBodySize (default 65536 bytes). Over the cap, an unknown content-encoding, or a request whose body never completed capture, gets no GraphQL enrichment at all - not an error, not a parsing-log failure.

A matching user webStreams.reqTemplates rule always wins: GraphQL only fills the template when no user rule matched the request.

Request fields (req.graphql.*)​

FieldContent
operationTypequery / mutation / subscription, or batch for a batched request; empty for a hash-only persisted query
operationNameSelected operation name: the body's operationName, else the document's single named operation, else empty
rootFieldsTop-level selection field names of the selected operation, sorted, deduped, capped at 20
queryHashNormalised document hash (literals to placeholders, whitespace/comments stripped, aliases dropped, arguments/selections sorted), first 16 hex characters
persistedQueryHashThe full Automatic Persisted Query (APQ) sha256Hash, when the request carries one
variableNamesNames of the declared variables of the selected operation - never their values
batchSizeNumber of operations in the request (1 when not batched)
operationsBatch only: one identity string per member, capped at 20
statusRequest-side only value: INVALID_DOCUMENT when the body matched a GraphQL shape but did not parse, or parsed with no operation to select

The operation template (visible as req.template, and in the grid) takes one of these forms:

CaseTemplate
Named operationmutation CreateOrder
Anonymous operationquery order,customer (root fields, comma-joined, capped)
Persisted, operationName sentpersisted GetOrder
Persisted, no namepersisted ecf4a1b2 (first 8 hex characters of the hash)
Batchbatch[3] query GetA + mutation SetB + … (first members, then + …)
INVALID_DOCUMENTUnchanged - the HTTP template still applies

Response fields (res.graphql.*)​

Parsed from the JSON response body, under the same decode and size rules, whatever the HTTP status

  • servers disagree on it: a validation error can come back 200, 400 or 404 depending on the GraphQL server. Neither stats.statusCode nor the communication's own status is ever changed by this.
ValueMeaning
OKNo errors in the response
PARTIALerrors present, and at least one non-null root field in data
ERRORerrors present, and data is null, absent, or every root field in it is null
UNKNOWNThe body could not be parsed as a GraphQL result, or was over the size cap

Alongside status: errorCount (total errors[] length, summed across a batch response), errorCode (the first errors[].extensions.code that carries one, e.g. UNAUTHENTICATED), and errorMessage (the first error message that is a string - an error whose message is missing or not a string is counted in errorCount but does not fill errorMessage), capped at 256 bytes after masking.

Variables​

Only variable names ever reach req.graphql.variableNames or a template - values stay in the request body, behind the same content permission as any other payload. errorMessage goes through a value-marker mask: GraphQL servers routinely echo the offending value next to a fixed phrase (got invalid value ..., ... value: ..., with value '...'), so everything from that phrase to the end of the message is replaced by …, quoted or bare. Outside such a phrase, a GraphQL name, a variable reference ("$email") or a type reference ("String!") stays readable; any other quoted text is masked. errorMessage is not stored at all when the whisperer does not save content for that communication (saveContent, or saveTlsDecryptedContent for decrypted TLS) or its URI matches urisToFilterResContent; status, errorCount and errorCode are still stored. The one residual is a custom resolver's own message that echoes a bare value with none of those marker phrases - it is not masked, so treat errorMessage as untrusted for anything sensitive and prefer errorCode for grouping/aggregation.

GET caveat

The URL query string (req.query) is stored exactly as it is for any other GET request, and for a GraphQL GET it contains the operation's variable values in clear. This is existing HTTP behaviour, not something GraphQL parsing changes - restrict content/query visibility the same way you would for any other GET request carrying sensitive parameters.

Persisted queries​

A hash-only request (Automatic Persisted Queries) carries no document, so it has no operationType and no rootFields to report - the platform never guesses them from a prior request. Its template is persisted <name> when the client sent an operationName, else persisted <hash8> (the first 8 hex characters of persistedQueryHash). There is no name-learning across requests: two hash-only calls to the same persisted query, one with a name and one without, produce two different templates.

Subscriptions​

GraphQL subscriptions carried over WebSocket or SSE are enriched the same way as request/response GraphQL, one document per message of the long-lived stream (see the WebMessageStream reference for the message schema itself).

  • WebSocket - the connection is GraphQL when its negotiated subprotocol is graphql-transport-ws (graphql-ws) or graphql-ws (legacy subscriptions-transport-ws), on any path. With no subprotocol negotiated, the connection is only enriched on a path configured in server.graphql.paths (or under heuristic).
  • SSE, distinct-connections and native - detected the same way as an ordinary GraphQL request (see Detected shapes above): the opening request itself is a GraphQL request (POST JSON, POST raw, or GET with query=).
  • SSE, single-connection mode - the stream request carries no GraphQL document of its own (the client POSTs subscribe/complete operations on a separate connection); it is recognised only by its path being configured in server.graphql.paths (or under heuristic). See the caution below for what this costs.

Each message of the operation carries req.graphql.* identity - operationType, operationName, rootFields, queryHash, persistedQueryHash - on both directions, with one exception: graphql-sse single-connection mode never has req.graphql on any message, by design - see the caution below, which covers why and what to do about it instead of relying on req.graphql there. The subscribe/start message that opens the operation additionally carries variableNames (and batchSize: 1); later messages of the same operation do not repeat it. res.graphql.* outcome - status, errorCount, errorCode, errorMessage - is set only on server-to-client result/error messages (WebSocket next/data/error, SSE next). The split for the rest is binary: complete/stop messages always carry req.graphql identity and the operation template, never an outcome (again, except single-connection SSE, which has no identity to carry on any message type); connection_init/ connection_ack, ping/pong and ka messages carry no GraphQL field at all.

The operation template (req.template) is inherited by every message of the operation from its handshake, but only when the handshake itself has no user template (empty, or -) - a user template set on the handshake (a matching webStreams.reqTemplates rule, see above) wins instead, exactly as for an ordinary GraphQL request. For WebSocket, server.webSocket.reqTemplates rules still apply on top of the inherited value, using the $t token to reference it: a rule named $t / inbound on the inherited template subscription P9_WatchA yields subscription P9_WatchA / inbound. For SSE, the handshake already carries the final operation template, so its events simply inherit it the usual way.

graphql-sse single-connection mode

When a graphql-sse deployment separates the stream from the operation - the client opens the SSE stream on one connection and posts subscribe/complete operations on another - the stream's next events carry res.graphql outcome only, with no req.graphql identity to pair it with, since the POST that started the operation is a separate HTTP exchange the stream never sees. Native graphql-sse, where the subscribe request is the SSE handshake itself, does not have this gap.

Only the SSE next and complete event types are read. A result sent under the default message event type (no event: field naming next) is not enriched.

Not parsed​

  • Multipart uploads (multipart/form-data, the GraphQL multipart request spec) - not detected.
  • Bodies over server.graphql.maxBodySize - no enrichment, silently.