The Inferock Baseline

The requirements a model route must meet to qualify with Inferock, for its exact model, provider route, region and configuration.

Status: proposed requirements. No completed route assessment is supplied.

Proposed evaluation standard · inferock-baseline-v1 · September 23, 2026

Overview

These 38 requirements define how a model route is evaluated for quality, performance, privacy, billing and service. Targets describe the aim; minimums and mandatory gates determine qualification within the assessed scope.

Requirement categories

What an assessment covers

A route assessment applies to one exact model revision, provider route, region, client and policy, and workload and traffic envelope. Its results do not carry over to any other route or configuration.

The Inferock Baseline defines qualification requirements for a model route; a measurement baseline, as used in Methodology, is the reference result a new measurement is compared against.

  1. Customer-controlled client
  2. Encrypted content path
  3. Verified workload decrypts to compute
Intended architecture; qualification requires evidence for the exact route and configuration.

The intended encrypted path begins in a customer-controlled client and ends inside a verified confidential workload. The approved workload decrypts the content to run the model; the response takes the protected return path. Each route needs evidence for this complete path.

These requirements do not mean every model or connection is encrypted. Encryption and privacy claims for a verified confidential route do not extend to ordinary Inferock Watch connections or to any route without its own verified confidential workload and encrypted path.

Gate/applicability

  • Qualification applies to the assessed model revision, provider route, region, client and policy, workload and configuration. All applicable minimums must pass; an overall average cannot conceal a failed gate.
  • Five quality requirements and all 12 encryption and privacy requirements are mandatory gates. Passing other checks cannot compensate for a failed gate.
  • N/A requires a documented reason for a genuinely optional, unsupported check. A missing required capability disqualifies the route. Untested or missing evidence is not a pass.
  • Enterprise operations requirements apply to the contracted tier. They are proposed requirements, not a universal support or service commitment.

Quality · 19 requirements

Q01 Broken outputMinimum ≤0.10% invalidTarget ≤0.01% invalidGate/applicability Applicable minimum required

Definition

Completed outputs that violate a declared, supported schema / schema-bound eligible attempts. Target 99.99% validity. Valid refusals are separately labelled; truncation and service failures remain visible in their own rows.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.01% invalid supported-schema outputs
  • Qualification requirement — maximum allowed: 0.1% invalid supported-schema outputs

Minimum

≤0.10% invalid

Target

≤0.01% invalid

Gate/applicability

Applicable minimum required

Evidence required

Use constrained decoding and client/enclave validation; syntax is distinct from task correctness.

Q02 Unexpected truncationMinimum ≤0.20%Target ≤0.05%Gate/applicability Applicable minimum required

Definition

Provider-caused premature completions / eligible long-output attempts. Use a sufficient, declared output budget; exclude deliberate user stops and deliberately undersized caps, recording both separately.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.05% provider-caused premature completions
  • Qualification requirement — maximum allowed: 0.2% provider-caused premature completions

Minimum

≤0.20%

Target

≤0.05%

Gate/applicability

Applicable minimum required

Evidence required

First-attempt rate; do not conceal failures with retries.

Q03 Billed empty outputMinimum 0 unresolved eventsTarget 0 unresolved eventsGate/applicability Mandatory gate

Definition

Charged operations with no usable requested output, after reconciliation. Valid tool payloads count as output; an explicitly valid refusal is separately labelled. Record raw empty events and their charges too.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0 unresolved billed-empty-output events
  • Qualification requirement — maximum allowed: 0 unresolved billed-empty-output events

Minimum

0 unresolved events

Target

0 unresolved events

Gate/applicability

Mandatory gate

Evidence required

Billing-integrity gate. Zero tolerance is an operating rule, not a claim that a finite sample proves impossibility.

Q04 OpenAI token recountMinimum ≤0.50% review bandTarget ≤0.10% unexplained excessGate/applicability Applicable minimum required

Definition

Sum of positive unexplained input-token differences / comparable recounted input tokens. Match model tokenizer and message framing; separate known reasoning, multimodal and system components.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.1% positive unexplained input-token excess
  • Qualification requirement — maximum allowed: 0.5% positive unexplained input-token excess

Minimum

≤0.50% review band

Target

≤0.10% unexplained excess

Gate/applicability

Applicable minimum required

Evidence required

Only where this adapter can reconcile components. OpenAI-compatible syntax alone does not establish accurate local recounts. Review band is not an allowance to overbill.

Q05 Anthropic token cross-checkMinimum ≤1.00% review bandTarget ≤0.50% aggregate differenceGate/applicability Applicable minimum required

Definition

Absolute aggregate difference between comparable count-endpoint estimates and normalized input usage / estimated input tokens. Keep documented system additions and unsupported server-tool inputs separate.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.5% absolute aggregate input-token difference
  • Qualification requirement — maximum allowed: 1% absolute aggregate input-token difference

Minimum

≤1.00% review band

Target

≤0.50% aggregate difference

Gate/applicability

Applicable minimum required

Evidence required

A calibration target, not proof of overcharging. Native Anthropic checks can be N/A on a non-Anthropic Inferock Serve route.

Q06 Request identity / duplicate chargingMinimum 0 unexplained collisionsTarget 0 unexplained collisionsGate/applicability Mandatory gate

Definition

Unexpected reuse of an identity for distinct logical requests, or a duplicate charge for one idempotent operation. Legitimate retries sharing an idempotency key are not automatically failures.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0 unexplained request-identity collisions or duplicate-charge events
  • Qualification requirement — maximum allowed: 0 unexplained request-identity collisions or duplicate-charge events

Minimum

0 unexplained collisions

Target

0 unexplained collisions

Gate/applicability

Mandatory gate

Evidence required

Hard gate: canonical operation ID, attempt IDs and at-most-once ledger effects.

Q07 Cache pricing anomalyMinimum ≤0.50% review bandTarget ≤0.10% unreconciled deviationGate/applicability Applicable minimum required

Definition

Absolute unreconciled cache-charge difference / expected cache charges across reconcilable requests. Use documented cache reads/writes, price version, tiers and rounding.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.1% unreconciled cache-charge deviation
  • Qualification requirement — maximum allowed: 0.5% unreconciled cache-charge deviation

Minimum

≤0.50% review band

Target

≤0.10% unreconciled deviation

Gate/applicability

Applicable minimum required

Evidence required

Cache-disabled routes may be N/A. A documented full-price cache miss is not automatically an anomaly.

Q08 Availability, uptime and error rateMinimum ≥99.90% availableTarget ≥99.95% availableGate/applicability Applicable minimum required

Definition

Eligible operations receiving a complete protocol-valid response within their maximum deadline / all eligible operations, before retries. Count provider 5xx, in-envelope 429, timeouts and broken streams; valid safety refusals count as available service. Fleet view: retain numerator/denominator per model, provider, region and traffic slice. Report time-based endpoint uptime separately; do not equate it to successful-request availability or average away a failing route.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.05% eligible operations without a complete protocol-valid response by their deadline
  • Qualification requirement — maximum allowed: 0.1% eligible operations without a complete protocol-valid response by their deadline

Minimum

≥99.90% available

Target

≥99.95% available

Gate/applicability

Applicable minimum required

Evidence required

Monthly request SLI plus 7-day rolling view. At one million requests, the floor allows 1,000 failures; the reference target allows 500. This is not a minutes-of-uptime measure. Deduplicates Quality Variations “Provider downtime” and “Uptime & error rate”.

Q09 Latency distribution and generation speedMinimum Chat: TTFT p95 ≤2 s / p99 ≤5 s; TPOT p95 ≤40 ms; completion p95 ≤13 sTarget Chat: TTFT p95 ≤1 s / p99 ≤3 s; TPOT p95 ≤20 ms; completion p95 ≤7 sGate/applicability Applicable minimum required

Definition

Client-observed first usable token on interactive non-reasoning 2k-input / 256-output traffic, established verified session, same-region client. Include routing and queue time. Use the separate workload profiles below. Record p50/p95/p99 separately for first token, token spacing and complete response, without averaging percentiles across providers. Qualification uses the specified p95/p99 limits; other quantiles are reported distributions. Long context and reasoning have their own profiles below. An unexpected cold start on the admitted warm tier remains in that tier’s latency/failure accounting.

How measured

Requirements (not observed results)

Interactive workload — first-token summary (95th percentile)

Established verified session; non-reasoning traffic; same-region client.

  • Target requirement — maximum allowed: 1 second to first usable token (95th percentile)
  • Qualification requirement — maximum allowed: 2 seconds to first usable token (95th percentile)
Target requirementQualification requirement
Interactive workload — non-reasoning, established verified session, same-region client
Test workload input size2000 tokens
Test workload output size256 tokens
Maximum allowed time to first usable token (95th percentile)1 second2 seconds
Maximum allowed time to first usable token (99th percentile)3 seconds5 seconds
Maximum allowed time per output token (95th percentile)20 milliseconds40 milliseconds
Maximum allowed response-completion time (95th percentile)7 seconds13 seconds
Maximum allowed response deadline30 seconds
Long-context workload
Test workload input size32000 tokens
Test workload output size1000 tokens
Maximum allowed time to first usable token (95th percentile)4 seconds8 seconds
Maximum allowed time to first usable token (99th percentile)10 seconds15 seconds
Maximum allowed time per output token (95th percentile)40 milliseconds67 milliseconds
Maximum allowed response-completion time (95th percentile)45 seconds80 seconds
Maximum allowed response deadline180 seconds
Reasoning workload
Test workload input size8000 tokens
Test workload reasoning budget2000 tokens
Test workload visible output size1000 tokens
Maximum allowed time to first visible token (95th percentile)30 seconds60 seconds
Maximum allowed time per visible output token (95th percentile)40 milliseconds67 milliseconds
Maximum allowed response-completion time (95th percentile)90 seconds150 seconds
Maximum allowed response deadline300 seconds

Fresh verification/setup — fresh session

  • Target requirement — maximum fresh verification/setup time (95th percentile): 3 seconds
  • Qualification requirement — maximum fresh verification/setup time (95th percentile): 5 seconds

Minimum

Chat: TTFT p95 ≤2 s / p99 ≤5 s; TPOT p95 ≤40 ms; completion p95 ≤13 s

Target

Chat: TTFT p95 ≤1 s / p99 ≤3 s; TPOT p95 ≤20 ms; completion p95 ≤7 s

Gate/applicability

Applicable minimum required

Evidence required

Companion targets: p95 per-request TPOT ≤20 ms (reference) / ≤40 ms (floor), about 50 / 25 tokens/s at the slow end. Report total completion and cold-session delay separately. Fresh verification/setup p95 ≤3 s target / ≤5 s ceiling; measure separately and include it in fresh-session total latency. Client network, routing and queue delay remain visible. Fleet latency percentiles map here, not to a duplicate row.

Q10 Unnecessary refusalMinimum ≤0.50% benign refusalsTarget ≤0.10% benign refusalsGate/applicability Applicable minimum required

Definition

Unjustified refusals / pre-labelled benign, supported prompts. Correct safety refusals on disallowed prompts are assessed separately and must not be optimized away.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.1% unjustified refusals on supported benign prompts
  • Qualification requirement — maximum allowed: 0.5% unjustified refusals on supported benign prompts

Minimum

≤0.50% benign refusals

Target

≤0.10% benign refusals

Gate/applicability

Applicable minimum required

Evidence required

Use a task-, language- and policy-matched corpus. These are proposed thresholds, not a universal safety benchmark.

Q11 Quality retained after changesMinimum ≥99.0% reference retainedTarget ≥99.5% reference retainedGate/applicability Applicable minimum required

Definition

100 × max(0, 1 − new quality score / fixed same-model reference score), on identical versioned tasks and sampling settings. The reference is non-zero and high precision; compare uncertainty intervals too.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.5% relative quality-score drop
  • Qualification requirement — maximum allowed: 1% relative quality-score drop

Minimum

≥99.0% reference retained

Target

≥99.5% reference retained

Gate/applicability

Applicable minimum required

Evidence required

Freeze the reference before testing. Do not relabel a provider’s degraded quantization as its own new baseline.

Q12 Tool-call validityMinimum ≥99.90% structurally validTarget ≥99.99% structurally validGate/applicability Applicable minimum required

Definition

Malformed tool names/arguments or violations of the declared supported schema / tool-call-required attempts. Semantic tool selection gets a separate task-correctness score.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.01% structurally invalid tool calls
  • Qualification requirement — maximum allowed: 0.1% structurally invalid tool calls

Minimum

≥99.90% structurally valid

Target

≥99.99% structurally valid

Gate/applicability

Applicable minimum required

Evidence required

Only required/supported tool workflows; an unsupported mandatory feature disqualifies the route instead of becoming N/A.

Q13 Security / governance signalsMinimum 0 critical exposuresTarget 0 critical exposuresGate/applicability Mandatory gate

Definition

Zero confirmed canary-secret leakage or forbidden governance outcomes in the authorized bounded evaluation; require 100% configured policy/evidence fields present.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0 critical events in the authorized bounded evaluation
  • Qualification requirement — maximum allowed: 0 critical events in the authorized bounded evaluation

Minimum

0 critical exposures

Target

0 critical exposures

Gate/applicability

Mandatory gate

Evidence required

This canonical detector row does not certify TEE/E2EE.

Q14 Grounded factual accuracyMinimum ≥95% correctTarget ≥98% correctGate/applicability Applicable minimum required

Definition

Incorrect answers / all questions in a fixed, answerable, source-grounded corpus with a predeclared rubric. Count unjustified abstentions as misses and report coverage; keep unanswerable safety/abstention tests separate.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 2% incorrect answers in the fixed source-grounded corpus
  • Qualification requirement — maximum allowed: 5% incorrect answers in the fixed source-grounded corpus

Minimum

≥95% correct

Target

≥98% correct

Gate/applicability

Applicable minimum required

Evidence required

Our proposed task threshold. Not a 98% truthfulness claim for arbitrary prompts, legal advice, medicine or every model.

Q15 Unexpected filtered omissionMinimum ≤0.10% benign omissionsTarget ≤0.05% benign omissionsGate/applicability Applicable minimum required

Definition

Unexpected content-filter output removal / eligible benign requests where the adapter exposes filter evidence. Intentionally blocked unsafe content is not counted as a defect.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.05% unexpected content-filter omissions on eligible benign requests
  • Qualification requirement — maximum allowed: 0.1% unexpected content-filter omissions on eligible benign requests

Minimum

≤0.10% benign omissions

Target

≤0.05% benign omissions

Gate/applicability

Applicable minimum required

Evidence required

Provider-specific detection applies only to equivalent documented metadata; unexplained missing metadata is unknown, not a pass.

Q16 Clean stream completionMinimum ≥99.90% cleanly completeTarget ≥99.99% cleanly completeGate/applicability Applicable minimum required

Definition

Streams ending without the documented terminal marker/finish evidence / eligible started streams. Client cancellations are separate. Any final usage evidence required by the API must reconcile.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.01% streams missing required terminal or finish evidence
  • Qualification requirement — maximum allowed: 0.1% streams missing required terminal or finish evidence

Minimum

≥99.90% cleanly complete

Target

≥99.99% cleanly complete

Gate/applicability

Applicable minimum required

Evidence required

No retry or failover may splice two model outputs into an apparently single successful stream.

Q17 Retry amplificationMinimum ≤1.020 attempts/opTarget ≤1.005 attempts/opGate/applicability Applicable minimum required

Definition

Total upstream attempts / logical eligible operations, including failed attempts; the score uses excess above 1. Also report extra token spend and post-retry completion separately.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0.005 extra upstream attempts per logical eligible operation
  • Qualification requirement — maximum allowed: 0.02 extra upstream attempts per logical eligible operation

Minimum

≤1.020 attempts/op

Target

≤1.005 attempts/op

Gate/applicability

Applicable minimum required

Evidence required

Normal traffic, excluding declared controlled probes. At most one automatic retry for safe pre-output/idempotent calls inside the original deadline; no automatic replay of side-effecting tools.

Q18 Served-model identityMinimum 0 unauthorized mismatchesTarget 0 unauthorized mismatchesGate/applicability Mandatory gate

Definition

Requested pinned model revision versus returned identity and, where available, attested model manifest. Every routing decision must satisfy the approved model/configuration contract.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0 unauthorized served-model mismatch events
  • Qualification requirement — maximum allowed: 0 unauthorized served-model mismatch events

Minimum

0 unauthorized mismatches

Target

0 unauthorized mismatches

Gate/applicability

Mandatory gate

Evidence required

Identity gate. Matching a response label alone does not prove exact weights. No silent model or privacy downgrade.

Q19 Known price before routingMinimum 0 unpriced billable routesTarget 0 unpriced billable routesGate/applicability Mandatory gate

Definition

Billable operations without an effective, versioned input/output/cache price and known service fee. Require quote coverage for 100% of billable routing decisions.

How measured

Requirements (not observed results)

  • Target requirement — maximum allowed: 0 unpriced billable-route events
  • Qualification requirement — maximum allowed: 0 unpriced billable-route events

Minimum

0 unpriced billable routes

Target

0 unpriced billable routes

Gate/applicability

Mandatory gate

Evidence required

Pricing gate. An absent price must not become a zero-cost result.

Fleet · 5 requirements

F01 Capacity, throughput and headroomMinimum ≥25% capacity above admitted peak; CC throughput loss ≤15%Target Certified capacity ≥1.25× forecast/admitted peak; CC throughput loss ≤10%Gate/applicability Applicable minimum required

Definition

Record sustainable successful output tokens/s, RPM, input TPM, output TPM and concurrency together, at Q08/Q09 quality/service limits and the required privacy profile. Qualify each model/provider/region and then the aggregate fleet. C ≥1.25× demand means at least 20% of certified capacity is spare, not 25% of capacity. Do not count unreserved supplier burst capacity as committed capacity.

How measured

Requirements (not observed results)

Capacity relative to admitted peak demand

  • Minimum required certified capacity relative to admitted peak demand: 1.25 × admitted peak demand
  • Maximum allowed admitted share of certified capacity: 0.8 of certified capacity

Confidential computing overhead — matched on/off configurations

  • Target requirement — maximum throughput loss with confidential computing: 10%
  • Qualification requirement — maximum throughput loss with confidential computing: 15%
  • Target requirement — maximum increase in time per output token with confidential computing: 10%
  • Qualification requirement — maximum increase in time per output token with confidential computing: 15%

Minimum

≥25% capacity above admitted peak; CC throughput loss ≤15%

Target

Certified capacity ≥1.25× forecast/admitted peak; CC throughput loss ≤10%

Gate/applicability

Applicable minimum required

Evidence required

Use the full offered-load curve, burst tests and admitted traffic mix. Compare matched CC-on/off throughput and TPOT; TPOT overhead target ≤10%, ceiling ≤15%. Example: 700 peak RPM needs ≥875 tested RPM capacity (rounded planning value 900). To promise 900 admitted RPM with this headroom, qualify ≥1,125 RPM. Preserve tail latency and quality at that point.

F02 Cache hit rateMinimum ≥90% of attainable token hit rateTarget ≥95% of attainable token hit rateGate/applicability Applicable minimum required

Definition

For a declared warmed workload, let A be the attainable cached-token share after accounting for documented minimum prefix size, TTL, routing and cache capacity. Raw hit rate H = actually reused prompt tokens / all prompt tokens. Target H/A ≥0.95; floor 0.90. Example A=50%: H ≥47.5% target / ≥45% minimum. Report raw H and A alongside the normalized score.

How measured

Requirements (not observed results)

Declared warmed workload

  • Target requirement — minimum token hit rate relative to attainable token hit rate: 0.95 × attainable token hit rate
  • Qualification requirement — minimum token hit rate relative to attainable token hit rate: 0.9 × attainable token hit rate

Illustrative example — attainable cached-token share of 50%

  • Example attainable cached-token share: 50%
  • Example target requirement — minimum raw token hit rate: 47.5%
  • Example qualification requirement — minimum raw token hit rate: 45%

Minimum

≥90% of attainable token hit rate

Target

≥95% of attainable token hit rate

Gate/applicability

Applicable minimum required

Evidence required

Use exact-prefix fixtures and provider/runtime cache evidence, not merely a repeated prompt. Separate cache computation from billed discount coverage (Q07). Only customer-authorized, tenant-isolated retained-cache profiles qualify; default zero-retention routes mark this N/A with a reason. A=0 is N/A, never a fabricated 100% pass.

F03 Missed-cache costMinimum ≤5% of attainable cache savings lostTarget ≤2% of attainable cache savings lostGate/applicability Applicable minimum required

Definition

Avoidable missed savings / attainable savings on cache-eligible traffic with known versioned prices. Weight by the actual input/cache price difference, so this is a cost metric distinct from F02’s token ratio. Record the dollar amount too. Exclude declared first fills and normal TTL expiry from avoidable misses; do not redefine the workload after measuring.

How measured

Requirements (not observed results)

  • Target requirement — maximum attainable cache savings lost: 2%
  • Qualification requirement — maximum attainable cache savings lost: 5%

Minimum

≤5% of attainable cache savings lost

Target

≤2% of attainable cache savings lost

Gate/applicability

Applicable minimum required

Evidence required

Replay a fixed eligible workload and reconcile cache state, routing decisions and billing evidence. If there is no supported paid discount, no savings opportunity or no authorized cache profile, mark this financial row N/A and report F02 where applicable. Zero-retention policy takes priority over a cache savings target.

F04 Rate-limit ownership and rejection rateMinimum 100% attribution; in-envelope 429s ≤0.05%Target 100% attribution; in-envelope 429s ≤0.01%Gate/applicability Applicable minimum required

Definition

Attribute each 429 to customer quota, Inferock admission or upstream/provider enforcement using route/attempt evidence. Unknown ownership remains unknown and fails the attribution target. In-envelope 429 rate = throttled eligible operations / all eligible operations inside the jointly agreed RPM/TPM/concurrency envelope. These errors also consume Q08’s availability budget; there is no extra excluded budget.

How measured

Requirements (not observed results)

  • Required attribution coverage for rate-limit rejections: 100%
  • Target requirement — maximum in-envelope rate-limit rejection rate: 0.01%
  • Qualification requirement — maximum in-envelope rate-limit rejection rate: 0.05%

Minimum

100% attribution; in-envelope 429s ≤0.05%

Target

100% attribution; in-envelope 429s ≤0.01%

Gate/applicability

Applicable minimum required

Evidence required

Exercise each layer’s quota independently and then within the committed combined envelope. Use content-free decision/attempt receipts, timestamps and documented limits. Provider 429s must not be relabeled customer misuse when the customer stayed inside the agreed contract.

F05 Answer consistency / contradictionMinimum ≤2% contradictory answer pairsTarget ≤1% contradictory answer pairsGate/applicability Applicable minimum required

Definition

On fixed, answerable factual tasks, proportion of paired answers that make incompatible material claims under identical evidence, model revision, prompt and sampling settings. Evaluate ≥1,000 independent tasks per required domain with declared repetitions; report uncertainty clustered by task. Creative variation and updated evidence are not contradictions.

How measured

Requirements (not observed results)

  • Target requirement — maximum contradictory answer-pair rate: 1%
  • Qualification requirement — maximum contradictory answer-pair rate: 2%
  • Minimum required independent tasks per domain: 1000 tasks per domain

Minimum

≤2% contradictory answer pairs

Target

≤1% contradictory answer pairs

Gate/applicability

Applicable minimum required

Evidence required

Use an adjudicated contradiction rubric and content-authorized synthetic fixtures or evaluation inside the trusted client/workload. Two consistently wrong answers can pass consistency and still fail factuality Q14. Do not log customer chats outside E01/E10 to calculate this metric.

Encryption and privacy · 12 requirements

E01 Where encryption begins and endsMinimum Mandatory: every stated condition must passTarget 100% content protected · 0 router decryptionGate/applicability Mandatory gate

Definition

Prompts, responses and any supported attachments, tool content and history are encrypted in the customer-controlled client and decrypted only in the approved confidential workload. The return path has the same protection. Ordinary Inferock or supplier routing services receive no plaintext. Metadata fields are explicitly enumerated.

How measured

Requirements (not observed results)

  • Required coverage of content encryption: 100%
  • Maximum allowed untrusted plaintext termination points: 0 termination points

Minimum

Mandatory: every stated condition must pass

Target

100% content protected · 0 router decryption

Gate/applicability

Mandatory gate

Evidence required

Trace a synthetic canary through the actual client, load balancer, Inferock gateway, provider gateway, worker and response. Capture safe network traces and inspect the configured termination points. A normal HTTPS proxy that decrypts content fails this route requirement.

E02 Encryption algorithms and strengthMinimum Mandatory: every stated condition must passTarget TLS 1.3 · ≥128-bit classical strengthGate/applicability Mandatory gate

Definition

Use TLS 1.3 plus an attestation-bound encrypted content channel: either TLS ending inside the verified workload, or a reviewed application protocol through outer TLS. Prefer AES-256-GCM; accept AES-128-GCM or ChaCha20-Poly1305 in an approved profile. Require full 128-bit authentication tags, unique per-key nonces and protocol usage limits. Use ephemeral X25519/P-256 or an approved stronger exchange. Disable content in TLS 0-RTT. AES-256 alone does not mean 256-bit strength for the entire protocol.

How measured

Requirements (not observed results)

  • Required TLS version: 1.3
  • Minimum required classical strength: 128 bits
  • Preferred AES key size: 256 bits
  • Required authentication-tag size: 128 bits
  • Maximum allowed accepted nonce reuse: 0 instances
  • Maximum allowed content-bearing TLS early data (0-RTT) requests: 0 requests

Minimum

Mandatory: every stated condition must pass

Target

TLS 1.3 · ≥128-bit classical strength

Gate/applicability

Mandatory gate

Evidence required

Record negotiated suites, library/build and content-protocol profile. Reject downgrade, altered ciphertext, reused nonce/replayed operation and wrong recipient. Verify key generation and authenticated context binding; do not design custom cryptography.

E03 Fresh proof tied to the connectionMinimum Mandatory: every stated condition must passTarget 100% sessions checked · evidence ≤5 min oldGate/applicability Mandatory gate

Definition

Before the first content in each new or renewed trust session, validate the complete CPU/workload/GPU trust chain, approved policy and content-channel key binding. The strict v1 target is evidence age ≤300 seconds, measured by the client/verifier; use a fresh 256-bit challenge or a reviewed equivalently fresh binding. Where signed timestamps are used, allow ≤30 seconds clock skew and require a trusted time source. A delegated verifier is permitted only with a pinned, verified policy and auditable evidence chain; a startup-only check does not silently satisfy this freshness target.

How measured

Requirements (not observed results)

  • Required trust-session verification coverage: 100%
  • Maximum allowed evidence age at client/verifier validation: 300 seconds
  • Required fresh challenge size, unless a reviewed equivalently fresh binding is used: 256 bits
  • Maximum allowed clock skew when signed timestamps are used: 30 seconds
  • Maximum allowed accepted invalid or replayed evidence: 0 instances

Minimum

Mandatory: every stated condition must pass

Target

100% sessions checked · evidence ≤5 min old

Gate/applicability

Mandatory gate

Evidence required

Test expired, replayed, wrong-nonce, wrong-key, wrong-audience, wrong-workload and revoked evidence. Save the verifier result and chain. For a provider mesh, map evidence age at every link and document any gap against this strict target.

E04 Protection while the model runsMinimum Mandatory: every stated condition must passTarget 100% compute protected · 0 debug modesGate/applicability Mandatory gate

Definition

Use a supported production CPU TEE and, for GPU inference, confidential mode on every participating GPU with protected CPU–GPU and any inter-GPU transfer. Verify device identity, firmware/security versions, revocation and production debug state. Disallow development/profiling modes that weaken isolation. Plaintext exists inside approved computation; protected GPU memory is not a claim that every HBM cell is encrypted.

How measured

Requirements (not observed results)

  • Required coverage of protected compute devices: 100%
  • Maximum allowed accepted debug or development-tool modes: 0 modes
  • Maximum allowed unprotected content-transfer paths: 0 paths

Minimum

Mandatory: every stated condition must pass

Target

100% compute protected · 0 debug modes

Gate/applicability

Mandatory gate

Evidence required

Collect CPU and separate device attestation, GPU readiness/admission outcome and topology. Try a missing/untrusted GPU, disabled CC, debug image and unprotected transfer path; all must block content. GPU subchecks can be N/A only for a declared CPU-only route.

E05 Who can obtain keysMinimum Mandatory: every stated condition must passTarget 0 operator keys · fresh session ≤60 minGate/applicability Mandatory gate

Definition

No Inferock, supplier or cloud administrator can export usable content/session keys or unilaterally alter a key-release policy trusted by the customer. A verified key service may hold keys inside its approved boundary. Establish fresh ephemeral key material and reverify trust at least every 60 minutes, or sooner at protocol usage limits or a security event. TLS KeyUpdate alone is not a new ephemeral exchange. Erase retired session keys after in-flight work finishes; maximum 5-minute drain, then stop and clear. A future server-key compromise must not unlock recorded past sessions.

How measured

Requirements (not observed results)

  • Maximum allowed operator-exportable content keys: 0 keys
  • Maximum allowed session-key age: 60 minutes
  • Maximum allowed drain time for retired keys: 5 minutes
  • Maximum allowed successful unauthorized key releases: 0 key releases

Minimum

Mandatory: every stated condition must pass

Target

0 operator keys · fresh session ≤60 min

Gate/applicability

Mandatory gate

Evidence required

Map all key holders, generation, rotation, storage and destruction. Exercise authorized KMS/IAM override, key export, old-session decryption and cloned-workload tests. An HPKE label alone does not demonstrate forward secrecy with a long-lived recipient key.

E06 Content retention and cleanupMinimum Mandatory: every stated condition must passTarget 0 durable plaintext · cleanup ≤60 sGate/applicability Mandatory gate

Definition

Default: zero prompt/response retention after request completion and zero plaintext content or reusable content keys in disks, logs, traces, crash dumps, swap and backups. Actively clear per-request CPU/GPU buffers and KV cache within 60 seconds of completion/cancellation and before any cross-tenant reuse. Clear retired session secrets under E05. Optional retained history or prompt caching needs a separately declared, customer-authorized encrypted profile with TTL and tenant-bound access; it does not receive the default zero-retention tag.

How measured

Requirements (not observed results)

  • Maximum allowed durable plaintext content: 0 bytes
  • Maximum allowed saved content records after a request in the default profile: 0 records
  • Maximum allowed transient-content cleanup time after completion or cancellation: 60 seconds
  • Maximum allowed cross-tenant reuse before cleanup: 0 instances

Minimum

Mandatory: every stated condition must pass

Target

0 durable plaintext · cleanup ≤60 s

Gate/applicability

Mandatory gate

Evidence required

Use synthetic canaries through success, cancellation, OOM, crash, restart and device reuse. Inspect writable sinks and cleanup behavior. Report logical clearing and hardware reuse protections; a software timeout alone is not proof of physical erasure of all memory.

E07 Administrator accessMinimum Mandatory: every stated condition must passTarget 0 privileged content reads or decryptsGate/applicability Mandatory gate

Definition

Production policy must deny operator shell/debug, memory-dump, snapshot/restore, injected-agent and credential paths that reveal content. Include administrators of Inferock, the inference supplier and cloud infrastructure. Replacing the workload or verifier policy must not trick a pinned customer into releasing content. Bound the claim to documented hardware, client, verifier and software assumptions.

How measured

Requirements (not observed results)

  • Maximum allowed successful unauthorized privileged content reads: 0 reads
  • Maximum allowed successful unauthorized decryptions: 0 decryptions
  • Maximum allowed undisclosed privileged content-access paths: 0 paths

Minimum

Mandatory: every stated condition must pass

Target

0 privileged content reads or decrypts

Gate/applicability

Mandatory gate

Evidence required

Run an authorized administrator attack matrix, including guest management APIs, host snapshots, debug settings and downstream credentials. Require code/configuration review as well as failed probes. Hardware-vendor and malicious-client compromise remain explicit threat-model limits.

E08 Approved client, code and modelMinimum Mandatory: every stated condition must passTarget 100% artifacts pinned · 0 unapproved changesGate/applicability Mandatory gate

Definition

Bind approved digests/provenance for the client/verifier, boot/runtime policy, containers, model weights, tokenizer, adapters, prompt template and security-relevant configuration. Require an auditable build-to-measurement chain and customer-controlled trust/update policy. Reject unauthorized rollback. If weights load after boot, trusted code must verify their digests before use. A browser client served mutable JavaScript by the operator needs a separate client-integrity solution; an on-page green check alone is insufficient.

How measured

Requirements (not observed results)

  • Required coverage of approved, pinned security-relevant artifacts: 100%
  • Maximum allowed accepted unauthorized changes or rollbacks: 0 changes or rollbacks

Minimum

Mandatory: every stated condition must pass

Target

100% artifacts pinned · 0 unapproved changes

Gate/applicability

Mandatory gate

Evidence required

Change one artifact at a time, substitute model weights, replay an old release and alter the client/update manifest. The verifier or workload admission must reject each unauthorized change. Save versioned build provenance and review records.

E09 Separation between customersMinimum Mandatory: every stated condition must passTarget 0 leaks in ≥10,000 boundary probesGate/applicability Mandatory gate

Definition

Bind 100% of session keys, authorization, request context and caches to the correct tenant. No cross-tenant content or secret disclosure. Run at least 10,000 authorized probes across parallel requests, reused connections, cancellation, cache hits, process restarts and GPU memory reuse; include every supported path. A mandatory customer workflow cannot be marked N/A to pass.

How measured

Requirements (not observed results)

  • Required tenant-binding coverage: 100%
  • Minimum required authorized tenant-boundary probes: 10000 probes
  • Maximum allowed cross-tenant disclosures: 0 disclosures

Minimum

Mandatory: every stated condition must pass

Target

0 leaks in ≥10,000 boundary probes

Gate/applicability

Mandatory gate

Evidence required

Publish the probe matrix, attempted/succeeded counts, canary method, tenant configuration and independent review. Ten thousand negative probes are a coverage floor, not a probability or proof of absolute safety.

E10 No training and no content collectionMinimum Mandatory: every stated condition must passTarget 0 training paths · 0 content exportsGate/applicability Mandatory gate

Definition

The pinned workload must not run training, fine-tuning, persistent learning or content collection for later use. No optimizer/training writes or writable model updates. Allow content-bearing egress only along the encrypted, authorized response path; prevent content exfiltration through tools, errors, logs, telemetry and abuse-monitoring exports. Features that send content to tools need separate customer authorization and an explicitly extended trust boundary.

How measured

Requirements (not observed results)

  • Maximum allowed enabled training or persistent-learning paths: 0 paths
  • Maximum allowed training or model-update writes: 0 writes
  • Maximum allowed unauthorized content exports: 0 exports
  • Maximum allowed unapproved content-egress destinations: 0 destinations

Minimum

Mandatory: every stated condition must pass

Target

0 training paths · 0 content exports

Gate/applicability

Mandatory gate

Evidence required

Review code, dependencies, writable mounts, credentials and egress policy. Use canaries to test content export, telemetry, training triggers and covert content in nominal metadata fields. A no-training contract or unchanged weight hash alone cannot establish this full guarantee.

E11 Errors, retries and fallbackMinimum Mandatory: every stated condition must passTarget 0 privacy downgrades · 100% safe recoveryGate/applicability Mandatory gate

Definition

Every retry, reconnect, stream restart and alternative provider must satisfy the same required security policy and approved model contract. Re-attest and re-encrypt for the actual recipient. Do not silently switch to ordinary HTTPS or an unqualified model route. Stop safely when no qualifying route exists; never trade privacy for availability.

How measured

Requirements (not observed results)

  • Maximum allowed plaintext or unqualified fallbacks: 0 fallbacks
  • Required qualification coverage for recovery paths: 100%

Minimum

Mandatory: every stated condition must pass

Target

0 privacy downgrades · 100% safe recovery

Gate/applicability

Mandatory gate

Evidence required

Force certificate/attestation failure, timeout, upstream outage, quota denial, mid-stream disconnect and all-route failure. Record the selected route and client-verification result. Test supported streaming and tools explicitly; unsupported features remain visibly unsupported.

E12 Customer evidence and ongoing validityMinimum Mandatory: every stated condition must passTarget 100% proof-linked requests · revoke ≤5 minGate/applicability Mandatory gate

Definition

Each served request has a content-free, authenticated receipt linking it to the verified session, exact route/model/policy and review version. Customers can inspect/reproduce verification. Requalify on every material hardware, firmware, runtime, model, client or policy change; reassess the security dossier at least every 90 days. Propagate a known revocation or disqualification through the controlled routing system within 5 minutes of its ingestion, block new admissions and terminate affected sessions safely. Source-update detection delay must be measured separately.

How measured

Requirements (not observed results)

  • Required coverage of requests linked to a verifiable session: 100%
  • Required requalification coverage for material changes: 100%
  • Maximum allowed security-review age: 90 days
  • Maximum allowed revocation propagation time after ingestion: 5 minutes
  • Maximum allowed accepted expired qualifications: 0 qualifications

Minimum

Mandatory: every stated condition must pass

Target

100% proof-linked requests · revoke ≤5 min

Gate/applicability

Mandatory gate

Evidence required

Save hardware evidence, endorsements/revocation data, measurements, key binding, policy decisions, relevant timestamps, test scope and reviewer. Test revocation, proof tampering, missing receipt, stale review and update admission. A signed receipt binds a session; it does not prove every individual GPU instruction executed correctly.

Enterprise operations · 2 requirements

S01 Critical-incident supportMinimum Same proposed minimum; requires staffed 24/7 P0 coverageTarget P0 acknowledgement ≤30 min; updates ≤60 min; incident report ≤5 business daysGate/applicability Applicable minimum required

Definition

Applies to the proposed $10k+ priority serverless agreement. Acknowledgement is distinct from resolution. The incident report clock starts after recovery. These commitments supplement Q08/Q09; they do not replace performance or privacy qualification.

How measured

Requirements (not observed results)

Proposed $10k+ priority serverless agreement

  • Maximum allowed critical (P0) incident acknowledgement time: 30 minutes
  • Maximum allowed interval between incident updates: 60 minutes
  • Maximum allowed incident-report time after recovery: 5 business days
  • Required daily critical-incident coverage: 24 hours per day

Minimum

Same proposed minimum; requires staffed 24/7 P0 coverage

Target

P0 acknowledgement ≤30 min; updates ≤60 min; incident report ≤5 business days

Gate/applicability

Applicable minimum required

Evidence required

Name the on-call owner and escalation path; rehearse an incident, keep response/update timestamps and provide the agreed report. Qualify staffing and upstream escalation before activating the contract.

S02 Service reporting and planned changesMinimum Same cadence for the proposed enterprise agreementTarget 1 SLO dashboard update/day; 1 service review/month; ≥7 days planned-change noticeGate/applicability Applicable minimum required

Definition

Report the contracted model, region, traffic envelope and per-row result. Planned model changes require at least seven days notice and customer policy approval. Emergency security changes follow a separately disclosed process, preserve E08/E11 and are promptly communicated. The universal per-request proof requirement stays in E12.

How measured

Requirements (not observed results)

Proposed enterprise agreement

  • Required service-level objective (SLO) dashboard update cadence: 1 update per day
  • Required service-review cadence: 1 review per calendar month
  • Minimum required planned-change notice: 7 days

Minimum

Same cadence for the proposed enterprise agreement

Target

1 SLO dashboard update/day; 1 service review/month; ≥7 days planned-change notice

Gate/applicability

Applicable minimum required

Evidence required

Keep dated dashboard snapshots, review records, release notices and approved manifests. A contract or notice cannot authorize silently sending traffic to a route that fails the customer’s privacy policy.

Route evidence

No route assessment is supplied with this version of the Inferock Baseline. When an assessment is supplied, its results appear here with the exact scope, date and evidence they apply to.

A route assessment must identify its model revision, provider route, region, client and policy, workload and traffic envelope, with dated observations and evidence for each applicable requirement.

Provider documentation, test coverage and observed results are different kinds of evidence. A supported check is not a passed check.