Http Parsing
The following services are managing HTTP communications parsing.
Below is their default configuration.
webWrite
webWrite is the Helm values key for the web parser (web-parser, renamed from web-write /
web-streams-write). Since the HTTP/2 work, it parses HTTP/1.x, HTTP/2 and gRPC together - there is
no separate gRPC parser service. Its configuration reflects that: one parsingJob reading a single
merged queue, and a protocols block carrying one store/cache set per protocol it may output.
"@id": web-parser
"@type": ServerConfiguration
version: '0.1'
protocols:
http:
comStore:
node: http://elasticsearch:9200
indexPurge: spider-search-httpcom-upload
indexGet: spider-search-httpcom
getTimeout: PT10S
purgeTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
contentStore:
node: http://elasticsearch:9200
indexGet: spider-search-httpcomcontent
indexPurge: spider-active-httpcomcontent-upload-default
getTimeout: PT15S
purgeTimeout: PT15S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT15S
circuitThreshold: 0.8
comsCache:
server: redis
port: 6379
db: 2
timeOut: PT10S
ttl: PT25S
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
comsContentCache:
server: redis
port: 6379
db: 12
timeOut: PT10S
ttl: PT25S
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
grpc: # present only when the grpcParsing feature flip is on
comStore:
node: http://elasticsearch:9200
indexPurge: spider-search-grpccom-upload
indexGet: spider-search-grpccom
getTimeout: PT10S
purgeTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
contentStore:
node: http://elasticsearch:9200
indexPurge: spider-search-grpccomcontent-upload
indexGet: spider-search-grpccomcontent
getTimeout: PT10S
purgeTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
comsCache: # a distinct Redis instance (redis-grpc), not the HTTP cache's db
server: redis-grpc
port: 6379
db: 2
timeOut: PT10S
ttl: PT25S
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
comsContentCache:
server: redis-grpc
port: 6379
db: 12
timeOut: PT10S
ttl: PT25S
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
parsingLog:
store:
node: http://elasticsearch:9200
indexPurge: spider-search-httppers-upload
indexGet: spider-search-httppers
getTimeout: PT10S
purgeTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
cache:
server: redis
port: 6379
db: 3
timeOut: PT10S
ttl: PT2M
circuitDuration: PT30S
circuitThreshold: 1
compressed: false
parsingJob:
pollingDelay: PT1S
jobRunners: 1
uri: http://tcp-update/v1/parsing-jobs/web
label: "/tcp-update/parsing-jobs/web"
timeout: PT10S
parsingDelay: PT10S
queues: ["web", "http", "grpc"] # "grpc" dropped from the list when grpcParsing is off
purgeJob:
sizeLimitKB: 100
parsingStatusSynchro:
delay: PT5S
queue: statsSynchro
maxBucketSize: 100
tcpSessions:
uri: http://tcp-update/v1/tcp-sessions
label: "/tcp-update/tcp-sessions"
timeout: PT10S
sizeLimitKB: 100
compressed: true
packetsByIndex:
uri: http://pack-read/v1/packets/of/tcpsession
label: "/pack-read/packets/of/tcpsession"
timeout: PT10S
whisps:
get:
uri: http://whisp/v1/whisperers/{id}
label: "/whisp/whisperers"
timeout: PT10S
config:
uri: http://whisp/v1/whisperers/{id}/config?view=full
label: "/whisp/whisperers/config"
timeout: PT10S
There is one parsingLog (parsing state / tracking resource), shared across protocols - a TCP session
being parsed for web traffic has a single parsing log regardless of whether it turns out to carry HTTP
or gRPC.
Two related settings live in the shared common configuration, not in webWrite itself:
parserProtocols- the list of protocols with a dedicated parsing queue (drives queue naming, e.g. theweb/http/grpcqueues above).outputProtocols- the list of protocols the web parser may emit a communication for (http,grpc,websocket,sse). Bothsseandwebsockethave an engine and appear in data - see SSE and WebSocket message parsing settings.
webRead
"@id": web-streams-read
"@type": ServerConfiguration
version: '0.1'
httpComStore:
node: http://elasticsearch:9200
indexGet: spider-search-httpcom
getTimeout: PT15S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT15S
circuitThreshold: 0.8
httpComContentStore:
node: http://elasticsearch:9200
indexGet: spider-search-httpcomcontent
getTimeout: PT15S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT15S
circuitThreshold: 0.8
httpPersStore:
node: http://elasticsearch:9200
indexGet: spider-search-httppers
indexParsingStatus: spider-search-parsing-status-httppers
getTimeout: PT10S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
packets:
uri: http://pack-read/v1/packets/payload/tcp/?force=true
label: "/packets/payload/tcp"
timeout: PT10S
searchRequest:
sizeLimitKB: 100
whisps:
uri: http://whisp/v1/whisperers/{id}/config?view=server
label: "/whisp/whisperers/config"
timeout: PT10S
webHttpComPoller
"@id": web-httpcom-poller
"@type": ServerConfiguration
version: '0.1'
logField: httpCom
itemStore:
node: http://elasticsearch:9200
useDataStoragePolicies: true
useDistinctPolicyForUpload: true
streamedDataStoragePolicies:
- name: default
indexSave: spider-active-httpcom-streaming-default
uploadedDataStoragePolicies:
- name: default
indexSave: spider-active-httpcom-upload-default
saveTimeout: PT5S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
itemCache:
server: redis
port: 6379
db: 2
timeOut: PT2S
ttl: PT45S
circuitDuration: PT15S
circuitThreshold: 1
compressed: false
polling:
queue: httpComToSynchronize
queueType: SORTED_SET
scoreAttribute: _update
size: 175
jobRunners: 1
delay: PT2S
automaticESIds: false
updateCache:
keepInCache: false
removeFromCacheCondition: {}
noTTLCondition: {}
whisps:
config:
uri: http://whisp/v1/whisperers/{id}/config?view=full
label: "/whisp/whisperers/config"
timeout: PT10S
webHttpComContentPoller
"@id": web-httpcom-content-poller
"@type": ServerConfiguration
version: '0.1'
logField: httpCom
itemStore:
node: http://elasticsearch:9200
useDataStoragePolicies: true
useDistinctPolicyForUpload: true
streamedDataStoragePolicies:
- name: default
indexSave: spider-active-httpcomcontent-streaming-default
uploadedDataStoragePolicies:
- name: default
indexSave: spider-active-httpcomcontent-upload-default
saveTimeout: PT5S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
itemCache:
server: redis
port: 6379
db: 12
timeOut: PT2S
ttl: PT45S
circuitDuration: PT15S
circuitThreshold: 1
compressed: false
polling:
queue: httpComContentToSynchronize
queueType: SORTED_SET
scoreAttribute: _update
size: 175
jobRunners: 1
delay: PT2S
automaticESIds: false
updateCache:
keepInCache: false
removeFromCacheCondition: {}
noTTLCondition: {}
whisps:
config:
uri: http://whisp/v1/whisperers/{id}/config?view=full
label: "/whisp/whisperers/config"
timeout: PT10S
webHttpPersPoller
"@id": web-httppers-poller
"@type": ServerConfiguration
version: '0.1'
logField: httpPers
itemStore:
node: http://elasticsearch:9200
useDataStoragePolicies: true
useDistinctPolicyForUpload: true
streamedDataStoragePolicies:
- name: default
indexSave: spider-active-httppers-streaming-default
uploadedDataStoragePolicies:
- name: default
indexSave: spider-active-httppers-upload-default
saveTimeout: PT5S
connectTimeout: PT2S
connectRetryDelay: PT15S
connectRetryTimes: 15
circuitDuration: PT30S
circuitThreshold: 1
itemCache:
server: redis
port: 6379
db: 3
timeOut: PT2S
ttl: PT45S
longTtl: PT3M30S
circuitDuration: PT15S
circuitThreshold: 1
compressed: false
polling:
queue: httpPersToSynchronize
queueType: SORTED_SET_HMAP
scoreAttribute: first
size: 175
jobRunners: 1
delay: PT5S
automaticESIds: false
updateCache:
keepInCache: true
useLongTtlAsSafety: true
removeFromCacheCondition:
property: status
values:
- COMPLETED
- ERROR
noTTLCondition: {}
saveInESConditions:
- property: state
values:
- CLOSED
- CLOSE_WAIT
- ESTABLISHED
- LAST_ACK
whisps:
config:
uri: http://whisp/v1/whisperers/{id}/config?view=full
label: "/whisp/whisperers/config"
timeout: PT10S
GraphQL
GraphQL is not a separate parser - it is an enrichment the web parser applies to the HTTP/1.x and
HTTP/2 communications it already produces on the paths configured in server.graphql.paths (default
["/graphql"]). Off the configured paths, nothing is detected unless heuristic is turned on, which
widens detection at the risk of mislabelling a REST body that happens to carry a query string key -
it is off by default for that reason.
Detected shapes:
- POST JSON -
content-type: application/json, body an object with a stringquery, or an object carryingextensions.persistedQuery.sha256Hash(a persisted / APQ hash-only request), or a non-empty array of such objects (a batch). - POST raw -
content-type: application/graphql, body is the document itself. - GET - the query string carries
query=(and optionallyoperationName=,variables=,extensions=).
The body used is the decoded one (gzip/deflate/br undone), capped at server.graphql.maxBodySize
(default 65536 bytes). Over the cap, an unknown content-encoding, or a request whose body never
completed capture, gets no GraphQL enrichment at all - not an error, not a parsing-log failure.
A matching user webStreams.reqTemplates rule always wins: GraphQL only fills the template when no
user rule matched the request.
Request fields (req.graphql.*)
| Field | Content |
|---|---|
operationType | query / mutation / subscription, or batch for a batched request; empty for a hash-only persisted query |
operationName | Selected operation name: the body's operationName, else the document's single named operation, else empty |
rootFields | Top-level selection field names of the selected operation, sorted, deduped, capped at 20 |
queryHash | Normalised document hash (literals to placeholders, whitespace/comments stripped, aliases dropped, arguments/selections sorted), first 16 hex characters |
persistedQueryHash | The full Automatic Persisted Query (APQ) sha256Hash, when the request carries one |
variableNames | Names of the declared variables of the selected operation - never their values |
batchSize | Number of operations in the request (1 when not batched) |
operations | Batch only: one identity string per member, capped at 20 |
status | Request-side only value: INVALID_DOCUMENT when the body matched a GraphQL shape but did not parse, or parsed with no operation to select |
The operation template (visible as req.template, and in the grid) takes one of these forms:
| Case | Template |
|---|---|
| Named operation | mutation CreateOrder |
| Anonymous operation | query order,customer (root fields, comma-joined, capped) |
Persisted, operationName sent | persisted GetOrder |
| Persisted, no name | persisted ecf4a1b2 (first 8 hex characters of the hash) |
| Batch | batch[3] query GetA + mutation SetB + … (first members, then + …) |
INVALID_DOCUMENT | Unchanged - the HTTP template still applies |
Response fields (res.graphql.*)
Parsed from the JSON response body, under the same decode and size rules, whatever the HTTP status
- servers disagree on it: a validation error can come back
200,400or404depending on the GraphQL server. Neitherstats.statusCodenor the communication's own status is ever changed by this.
| Value | Meaning |
|---|---|
OK | No errors in the response |
PARTIAL | errors present, and at least one non-null root field in data |
ERROR | errors present, and data is null, absent, or every root field in it is null |
UNKNOWN | The body could not be parsed as a GraphQL result, or was over the size cap |
Alongside status: errorCount (total errors[] length, summed across a batch response), errorCode
(the first errors[].extensions.code that carries one, e.g. UNAUTHENTICATED), and errorMessage
(the first error message that is a string - an error whose message is missing or not a string is
counted in errorCount but does not fill errorMessage), capped at 256 bytes after masking.
Variables
Only variable names ever reach req.graphql.variableNames or a template - values stay in the
request body, behind the same content permission as any other payload. errorMessage goes through a
value-marker mask: GraphQL servers routinely echo the offending value next to a fixed phrase (got invalid value ..., ... value: ..., with value '...'), so everything from that phrase to the end
of the message is replaced by