---
title: "How to support a service-level performance claim in a bid"
description: "Prove what a historical service result measured, which events counted, how the clock ran, and whether the result is comparable with the service now offered."
canonical: "https://zephior.com/insights/evidence-service-level-performance-in-a-bid"
last-updated: 2026-09-04
---

# How to support a service-level performance claim in a bid

> Prove what a historical service result measured, which events counted, how the clock ran, and whether the result is comparable with the service now offered.

By [Tony Kim](https://zephior.com/authors/tony-kim). Published 2026-09-04; updated 2026-09-04. 20 minute read.

## Definition

A service-level performance evidence record proves what one historical result actually represents before that result is used in a bid. It fixes the service and customer boundary, indicator, success or failure event, eligible population, measurement period, clock rules, numerator, denominator, exclusions, missing records, source lineage, calculation, approval and differences from the service now offered. The record then gives the claim a bounded use decision. It does not convert past performance into a future promise or prove that a buyer's proposed service level can be delivered.

## Problem

Service figures look comparable long before they are comparable. A dashboard says 99.98 percent availability, but it watched a central application rather than the user journey. A support report says 94 percent within four hours, but its clock stopped whenever the ticket awaited customer information. A field-service team reports two-hour response, but the event recorded was dispatch acceptance rather than arrival. Planned maintenance, customer-caused incidents, test traffic, duplicate tickets and missing sites may disappear before the percentage reaches the proposal. An evaluator sees one precise number while the evidence describes a different service, population or clock.

## Point of view

Start with the buyer-facing sentence, then rebuild the measurement from event records rather than trusting the dashboard label. A service-level claim is a fraction or distribution over an explicitly eligible population. Name what entered that population, what counted as success, when the clock began and ended, which pauses were permitted, and which records were absent or corrected. Compare the historical service with the offered service dimension by dimension. If the old result remains useful but cannot prove the broader proposition, preserve it with its boundary. Narrow evidence is not bad evidence; undisclosed expansion is the problem.

## Write down what the percentage is being asked to prove

A service report and a bid sentence serve different purposes. The report may help one operations team manage one contract. The bid sentence asks an evaluator to infer capability for another service. Begin by recording the exact proposed wording, the question and scoring descriptor it answers, whether the figure is qualification evidence, past-performance evidence or contextual support, and what a reasonable evaluator would understand. “We maintained 99.98 percent availability” can imply an entire service, every user, continuous coverage and a representative period even when none of those nouns appears.

Preserve the buyer's own metric separately. Copy its label, formula, service period, operating window, threshold, measurement point, exclusions and evidence request from the current tender documents. Do not retrofit the historical system to the buyer's vocabulary. Two indicators called availability may count different time, and two four-hour response measures may stop at different events. Similar labels are an invitation to compare definitions, not proof that the values share a meaning.

The fictional Alderpoint Transit Systems Ltd is bidding to monitor passenger-information displays at 48 stations. The buyer asks for evidence of end-to-end availability over a complete year. Its clock covers every minute of the 24-hour service, begins when a synthetic passenger request cannot reach a display and ends after a successful retest. Alderpoint's proposal draft says its existing service achieved 99.982 percent availability in 2025. The source dashboard, however, measured only the central message broker. This is a useful result, but it does not yet answer the buyer's proposition.

**The first fields in a service-level evidence record**

| Field | Question to settle | Alderpoint entry |
| --- | --- | --- |
| Proposed claim | What exact words will be released? | 99.982% availability in 2025 |
| Buyer use | What will the evaluator use it to judge? | End-to-end availability for 48 stations |
| Historical object | What actually produced the value? | Central message broker on one contract |
| Immediate issue | Why can the number not travel unchanged? | Station displays and telecom links were outside measurement |

## Turn the label into an event population and a calculation

Identify the unit over which performance is judged. Availability can be time-based, request-based, user-journey based or asset based. Incident performance can count tickets, affected services, customers or distinct outages. Field response can count calls, dispatches, arrivals or restored assets. State the eligible population before looking at the result. Otherwise the analyst can unintentionally choose a denominator that flatters the service.

For a ratio, preserve both parts. If 9,800 of 10,000 eligible requests succeeded, the result is 98 percent only under the recorded success definition. Requests rejected upstream, synthetic tests, retries and malformed calls may or may not belong in the population. For time availability, specify total eligible time, observed downtime and any time removed from both numerator and denominator. Do not mix “downtime removed from the denominator” with “downtime counted as available.” They can produce the same headline while asserting different facts.

Aggregations change meaning. An annual ratio of all events weights busy months more heavily. The average of twelve monthly percentages gives each month equal weight. The average performance across sites gives a quiet station the same weight as a central hub. A percentile answers a distribution question; a mean answers another. Google's SRE guidance distinguishes an indicator from its objective and agreement, and warns that averages can hide the tail. That distinction is essential in a proposal: show which measurement was made before describing whether it was good.

- Name the unit of analysis and every eligibility condition.
- Define success and failure in observable terms.
- Keep numerator, denominator and exclusions separately recoverable.
- Record aggregation, weighting, segmentation and rounding.
- Distinguish an observed result from a target and a contractual promise.

## Make every start, stop, pause and missing observation visible

A response time has no meaning until its events are named. Receipt by a monitored mailbox, creation in the ticket system, human acknowledgement, valid classification, assignment, remote action, technician arrival, restoration and closure are different timestamps. Choose the pair specified by the historical agreement or measurement rule and preserve any transformation. If a portal imported emails every five minutes, the ticket creation time is not the actual receipt time. If a dispatcher corrected an arrival time after the visit, both values and the reason should remain inspectable.

Define the calendar. State timezone, daylight-saving treatment, service hours, holidays, grace intervals and how events crossing a boundary are handled. ISO 8601 provides an exchange representation for dates and time offsets; it does not prove that two clocks were synchronized or that a business-hours calendar was correct. Near-threshold cases deserve particular scrutiny. A three-minute clock drift can reverse the classification of a 30-minute response.

Now inventory exclusions and pauses. Planned maintenance may be permitted, capped or counted. Time awaiting customer access may pause one contract but not another. A downstream carrier failure may be supplier risk in the new service even if the historical contract excluded it. Preserve the rule that existed during the observation period, when it was approved and how each excluded event was classified. A late spreadsheet filter called “non-service” is not a policy merely because it has been used before.

**Clock questions for common service measures**

| Measure | Possible start | Possible end | Common ambiguity |
| --- | --- | --- | --- |
| Availability | First failed observation | Verified recovery | Component or user journey |
| First response | Request received | Meaningful human response | Automatic acknowledgement |
| Restoration | Service impairment begins | Service usable again | Workaround versus full repair |
| On-site attendance | Valid call accepted | Technician arrives | Dispatch acceptance versus arrival |
| Resolution | Eligible case opens | Agreed outcome delivered | Closure despite reopening |

## Keep the result connected to the events that produced it

A screenshot proves that a display showed a value at one moment. It rarely proves the underlying population or calculation. Preserve the controlled report, its version, author, approval and period, but also identify the event sources, extraction time, query or calculation logic, field definitions, service calendar and correction log. A reviewer should be able to travel from the proposal sentence to the result, from the result to its numerator and denominator, and from those totals to inspectable records or an authorized assurance statement.

Keep lineage across transformations. Raw monitor events may be deduplicated, joined to an asset register, classified against a maintenance calendar and aggregated by month before a chart is created. Record each step, the responsible system or person, input version and output identifier. W3C PROV offers a useful vocabulary of entities, activities and responsible agents; the point is not to force an ontology into a bid, but to retain who did what to which evidence and when.

Protect confidential operational data. The proposal may cite an approved aggregate, client reference or assurance summary while raw incident descriptions remain in a controlled repository. Record disclosure authority, anonymization, retention and reviewer access. If the buyer requests the evidence, follow the procurement channel and permission boundary. Do not expose customer names, vulnerabilities, user data or infrastructure details merely to make the percentage look inspectable.

**Minimum provenance for one released result**

| Layer | Retain | What it prevents |
| --- | --- | --- |
| Claim | Exact wording and every proposal location | Different meanings around one number |
| Result | Value, period, unit, aggregation and approval | A dashboard snapshot without context |
| Calculation | Inputs, filters, formula, code or workbook version | An irreproducible percentage |
| Events | Source identifiers, timestamps and retained corrections | Totals detached from observations |
| Authority | Service owner, disclosure owner and decision time | Accidental release or unowned interpretation |

## Review the events most likely to disappear from the percentage

Test the negative space. Compare the expected asset, user or ticket population with what the measurement system saw. A monitor covering 37 of 48 stations has not produced evidence for all 48, even if every observed station performed perfectly. Look for periods when collection stopped, sites entered or left service, identifiers changed, tickets were merged, incidents were recategorized or data was backfilled. Missingness is not automatically failure, but it prevents a confident result until its cause and effect are understood.

Review exclusions as a population, not only one by one. Count their frequency, duration and share of total demand. Separate predeclared contractual exclusions from data-quality removal, test traffic, duplicates, withdrawn requests and discretionary management adjustments. Sample the underlying records. If customer-caused delay forms 30 percent of elapsed time, that fact may matter greatly when the new contract uses a continuous clock. A mathematically valid historical exclusion can still make the result unsuitable for the offered service.

Inspect variation. Report the complete period and material segments before selecting an illustration. Show monthly performance, eligible volume and failures, then analyze sites, channels, severity or load where they affect comparability. Do not choose the best month as a “representative” case. Do not average away a missed contractual threshold. A result can support the statement that annual event-weighted performance was 99.9 percent while failing to support “we met 99.9 percent every month.”

- Reconcile monitored objects with the authoritative service inventory.
- Quantify missing intervals and the denominator they could affect.
- Retain original and corrected event values with reason and approver.
- Profile exclusions by rule, volume, duration and service segment.
- Test the full time series before selecting any example.

## Decide whether the old service can speak for the new one

Build a side-by-side map of the historic service and the offered service. Compare legal entity, customer type, users, locations, channels, volume and peak load, criticality, technology, service components, operating hours, support tiers, delivery partners, dependencies, metric formula, exclusion regime and observation period. Mark each dimension same, materially similar, different, unknown or not applicable. Do not collapse the map into a generic similarity score. One critical difference can defeat an otherwise close match.

A change does not always invalidate the evidence. A larger service may still use the same proven component, and an improved monitoring design may make a conservative historical claim useful. Explain the logical link. If the offered design adds endpoint monitoring and redundant carrier paths, the broker result can evidence that broker's history while the architecture and test evidence address the new boundary. It cannot be relabelled as the availability of a system that did not exist.

Treat entity and supplier-chain changes explicitly. Performance achieved by an affiliate, incumbent, consortium member or subcontractor belongs to the entity and service that performed it. State that relationship and whether the same people, system, process or supplier will perform the relevant part of the offer. Past performance can inform evaluation, as FAR 15.305 illustrates, but relevance, source, context and trend remain part of the assessment. Ownership of a corporate logo is not operational continuity.

**Use decisions for the evidence record**

| State | Meaning | Permitted action |
| --- | --- | --- |
| directly_comparable | Material dimensions and metric match | Release exact result with source boundary |
| comparable_with_limits | Useful similarities remain with stated differences | Release narrowed claim and limitations |
| recalculation_required | Source events may support the buyer definition | Recalculate before any numeric claim |
| source_gap | Required records or definitions are absent | Hold the claim and seek evidence |
| exclusion_unresolved | Removed events may alter the result | Resolve or present no result |
| not_representative | Selected period or segment cannot support normal performance | Use only as a named example |
| historical_context_only | True result does not evidence the offered proposition | Describe only its historic object |
| claim_removed | No safe and useful formulation remains | Remove every occurrence |

## Let the limitation travel with the number

Write the supported proposition before writing persuasive copy. Name the measured object, metric, period and relevant boundary in the sentence or its immediate evidence note. “The central message-distribution component recorded 99.982 percent time availability during calendar 2025 under the incumbent contract's measurement method” is narrower than Alderpoint's draft, but it is defensible. Then state why it is relevant and where it is not equivalent to the proposed end-to-end metric.

Do not bury a material limitation in a distant footnote. A qualifier must be close enough that an evaluator or later reuser cannot detach the percentage from it. Avoid “proven,” “consistently,” “across our services,” “industry-leading” and “SLA achieved” unless the record supports those additional propositions. If every month met a threshold, say so only after testing each month. If the value is an annual aggregate, call it an annual aggregate.

Route the record to the service owner, measurement owner, disclosure authority and bid approver. Contract and commercial reviewers decide whether the offer can promise a target; the historical evidence owner does not. Link every occurrence in the response, executive summary, chart and reference form to the same approved wording. Reopen when the tender definition, source report, calculation, offered scope, architecture, supplier chain, permission or observation period changes.

## A true 99.982 percent result becomes a narrower, useful statement

Alderpoint retains the 2025 broker report, monitor configuration, monthly source exports, maintenance calendar, calculation workbook and approval. The value is reproduced as 99.982 percent over the broker's eligible annual minutes after the incumbent contract's permitted maintenance treatment. Monthly results range from 99.941 to 100 percent. The records support the central component and full calendar year. They do not contain station display status, telecom reachability or passenger-request completion.

A second endpoint dataset begins on 1 April 2025 and covers 37 stations. Eleven stations have no comparable endpoint monitor. Carrier incidents are tagged inconsistently, and several missing intervals cannot be classified. The team does not merge that incomplete dataset with the broker figure. It assigns source_gap for end-to-end history and historical_context_only to the broker value in relation to the buyer's exact proposition. The 99.982 percent number remains available for the narrower component claim.

The released answer explains that the broker history evidences stability of one proposed component, identifies the narrower measurement boundary, and presents the new end-to-end monitoring design separately. It does not claim that Alderpoint previously delivered the buyer's 48-station metric. Operations, commercial and contract owners decide the future availability target using the complete offered design, resources, dependencies and price. The evidence record closes only after every copied percentage uses the approved sentence; it reopens if the buyer revises the metric or the solution boundary changes.

- Supported: the central broker's recorded 2025 result under its historic method.
- Unsupported: end-to-end availability across all station displays and links.
- Unresolved: endpoint history for eleven stations and unclassified gaps.
- Separate decision: the service level that the new offer may commit to.
- Reopen triggers: buyer metric, architecture, source correction, scope or permission change.

## Useful outcomes

- One claim resolves to one service-level evidence record and a reproducible calculation.
- The measured service, users, locations, channels, components and operating hours are explicit.
- Eligible events and time are separated from excluded, missing, duplicated and test records.
- Clock start, stop, pause, timezone, aggregation and rounding rules are recoverable.
- Every result retains its observation period, source version and extraction time.
- Differences between the historical service and the offered service remain visible.
- An isolated month, best site or central component cannot silently represent normal performance.
- Approved wording states what the evidence proves and the limitations material to evaluation.
- Historical evidence never becomes an unowned future service-level commitment.

## Workflow

1. **Freeze the proposed claim and its evaluation use.** Record the exact sentence, tender question, scoring use, buyer definition, offered service and consequence if the statement is wrong.
2. **Define the historical service boundary.** Identify the entity, contract, service, customer group, sites, channels, components, hours, suppliers and period that generated the result.
3. **Reconstruct the indicator.** State the unit of analysis, eligible population, success event, failure event, numerator, denominator, target, aggregation and rounding.
4. **Rebuild the clock and exclusions.** Trace start and stop events, pause rules, service calendars, timezones, maintenance, dependencies, duplicates, reopenings and missing observations.
5. **Verify the source trail and calculation.** Preserve raw events or controlled reports, extraction logic, versions, corrections and approvals, then independently reproduce the published value.
6. **Test representativeness.** Inspect performance by month, site, severity, channel and load so a selected slice or average cannot conceal unstable or absent service.
7. **Compare history with the offered service.** List every material change in scope, architecture, geography, demand, operating hours, subcontractors, metric and contractual treatment.
8. **Authorize bounded wording and monitoring.** Choose the supported use, secure disclosure and service-owner approval, link all occurrences, and set expiry and reopening events.

## Key decisions

- What exact proposition will the evaluator infer from the number?
- Does the source measure a user outcome, an end-to-end service, one component or a proxy?
- Which requests, tickets, minutes, assets or visits were eligible to enter the calculation?
- Which event started, paused, restarted and ended the measurement clock?
- Were exclusions stated before observation or applied later to improve the result?
- How much of the expected population is missing, unmonitored or manually corrected?
- Do monthly and segment results support the aggregate or reveal material variation?
- Is the historic service sufficiently similar to the entity, scope and operating model offered?
- Which limitations must travel with the result for an evaluator to interpret it correctly?
- Who may approve disclosure and who may authorize any future commitment?

## Risks

- A component uptime figure may be presented as end-to-end service availability.
- A yearly average may conceal failed months or sites with no monitoring.
- The denominator may omit failed transactions that never reached the monitored component.
- Pause and exclusion rules may remove the events users experienced as failure.
- A response claim may measure acknowledgement, assignment, dispatch, arrival or restoration interchangeably.
- Reopened or duplicated incidents may be removed inconsistently.
- Timestamps from unsynchronized systems may create false precision near a threshold.
- Manual corrections may be accepted without reason, authority or retained before-state.
- A successful pilot or best-performing client may be portrayed as ordinary operations.
- Historical evidence may be mistaken for approval to promise the same target in the new contract.

## Metrics

- service-level claims with complete evidence records
- calculations independently reproduced from controlled sources
- eligible population covered by usable measurement records
- exclusions with a stated rule, reason and authority
- manual corrections retaining original value and approval
- claims tested across time and material service segments
- historical-to-offered service differences resolved or disclosed
- proposal occurrences linked to the approved wording
- claims narrowed or removed before release

## Frequently asked questions

### Is a dashboard screenshot enough to support an SLA claim?

Usually not by itself. Preserve the dashboard version and capture time, but also retain the metric definition, population, event sources, filters, exclusions, calculation, observation period and approval needed to reproduce and interpret the result.

### Can we quote the best-performing month?

Only as a clearly identified month for a legitimate reason. Do not present it as normal or representative performance. Show the complete period and material variation when the evaluator is judging sustained delivery.

### Does 99.9 percent availability have a standard meaning?

No. The value depends on the service boundary, eligible time or requests, success definition, measurement point, maintenance rule, exclusions, aggregation and period. Reconstruct those terms for each source and tender.

### How should missing monitoring data be treated?

Identify the missing objects and intervals, investigate why they are absent and test the maximum effect on the result. Do not silently count missing time as available or remove it from the denominator without an authorized rule.

### Can customer-caused delays be excluded?

Only according to the applicable historic metric and evidence. Preserve the exact pause or exclusion rule and affected events. Then explain any difference if the buyer's metric assigns that dependency differently.

### What if the historic and buyer metrics use different definitions?

Do not relabel the old value. Recalculate from source events if the buyer definition can be applied reliably. Otherwise use a narrower contextual claim, disclose the difference or remove the number.

### Does a historical result prove we can promise the same SLA?

No. History is evidence about a defined past service. A future commitment also requires an authorized offered design, demand basis, resources, dependencies, price, measurement regime, remedies and contract approval.

### When must the evidence record be reopened?

Reopen it after a change to the tender metric, offered service, historical source, calculation, exclusion decision, monitoring coverage, supplier relationship, disclosure permission or approved wording.


## Primary sources

- [Procurement Act 2023, section 23: award criteria](https://www.legislation.gov.uk/ukpga/2023/54/section/23), UK Legislation
- [Guidance on assessing competitive tenders under the Procurement Act 2023](https://www.gov.uk/government/publications/procurement-act-2023-guidance-documents-procure-phase/assessing-competitive-tenders-html), UK Cabinet Office
- [Model Services Contract combined schedules, version 2.2A](https://www.gov.uk/government/publications/the-model-services-contract-schedules-england-wales), UK Cabinet Office and Government Legal Department
- [Directive 2014/24/EU on public procurement](https://eur-lex.europa.eu/eli/dir/2014/24/oj/eng), European Union
- [GWB section 127: award](https://www.gesetze-im-internet.de/gwb/__127.html), Federal Ministry of Justice and Federal Office of Justice
- [VgV section 58: award and award criteria](https://www.gesetze-im-internet.de/vgv_2016/__58.html), Federal Ministry of Justice and Federal Office of Justice
- [French Public Procurement Code Article R2152-7](https://www.legifrance.gouv.fr/codes/article_lc/LEGIARTI000045739587), Légifrance
- [FAR 15.305: proposal evaluation](https://www.acquisition.gov/far/15.305), Acquisition.gov
- [FAR Subpart 42.15: contractor performance information](https://www.acquisition.gov/far/subpart-42.15), Acquisition.gov
- [World Bank rated criteria guidance for buyers and suppliers](https://www.worldbank.org/ext/en/what-we-do/project-procurement/rated-criteria), World Bank Group
- [World Bank guidance on evaluating bids and proposals with rated criteria](https://thedocs.worldbank.org/en/doc/9dcb7971706bf29b2732779c39922b77-0290012025/original/Evaluating-Bids-and-Proposals-with-Rated-Criteria-Feb-4-2025.pdf), World Bank Group
- [ISO/IEC 20000-1:2018 service management system requirements](https://www.iso.org/standard/70636.html), International Organization for Standardization
- [ISO/IEC/IEEE 15939:2017 measurement process](https://www.iso.org/standard/71197.html), International Organization for Standardization
- [ISO 8601-1:2019 date and time representations](https://www.iso.org/standard/70907.html), International Organization for Standardization
- [Google SRE book chapter on service level objectives](https://sre.google/sre-book/service-level-objectives/), Google Site Reliability Engineering
- [Google SRE book chapter on monitoring distributed systems](https://sre.google/sre-book/monitoring-distributed-systems/), Google Site Reliability Engineering
- [Semantic conventions for HTTP metrics](https://opentelemetry.io/docs/specs/semconv/http/http-metrics/), OpenTelemetry
- [How to set performance metrics for a government service](https://www.gov.uk/service-manual/measuring-success/how-to-set-performance-metrics-for-your-service), UK Government Service Manual
- [PROV-O: the W3C provenance ontology](https://www.w3.org/TR/prov-o/), World Wide Web Consortium
- [W3C Data Quality Vocabulary](https://www.w3.org/TR/vocab-dqv/), World Wide Web Consortium


## Related articles

- [Answer an RFP service-level question without overpromising](https://zephior.com/insights/answer-an-rfp-service-level-question)
- [Does this evidence cover your proposal claim?](https://zephior.com/insights/verify-the-scope-of-proposal-evidence)
- [Answer an RFP scalability question with measured limits](https://zephior.com/insights/answer-an-rfp-scalability-question)
- [Write an executive summary the bid can prove](https://zephior.com/insights/align-an-executive-summary-with-bid-evidence)
