A service-level performance evidence record proves what one historical result actually represents before that result is used in a bid. It fixes the service and customer boundary, indicator, success or failure event, eligible population, measurement period, clock rules, numerator, denominator, exclusions, missing records, source lineage, calculation, approval and differences from the service now offered. The record then gives the claim a bounded use decision. It does not convert past performance into a future promise or prove that a buyer's proposed service level can be delivered.

Service figures look comparable long before they are comparable. A dashboard says 99.98 percent availability, but it watched a central application rather than the user journey. A support report says 94 percent within four hours, but its clock stopped whenever the ticket awaited customer information. A field-service team reports two-hour response, but the event recorded was dispatch acceptance rather than arrival. Planned maintenance, customer-caused incidents, test traffic, duplicate tickets and missing sites may disappear before the percentage reaches the proposal. An evaluator sees one precise number while the evidence describes a different service, population or clock.

Start with the buyer-facing sentence, then rebuild the measurement from event records rather than trusting the dashboard label. A service-level claim is a fraction or distribution over an explicitly eligible population. Name what entered that population, what counted as success, when the clock began and ended, which pauses were permitted, and which records were absent or corrected. Compare the historical service with the offered service dimension by dimension. If the old result remains useful but cannot prove the broader proposition, preserve it with its boundary. Narrow evidence is not bad evidence; undisclosed expansion is the problem.

Write down what the percentage is being asked to prove

A service report and a bid sentence serve different purposes. The report may help one operations team manage one contract. The bid sentence asks an evaluator to infer capability for another service. Begin by recording the exact proposed wording, the question and scoring descriptor it answers, whether the figure is qualification evidence, past-performance evidence or contextual support, and what a reasonable evaluator would understand. “We maintained 99.98 percent availability” can imply an entire service, every user, continuous coverage and a representative period even when none of those nouns appears.

Preserve the buyer's own metric separately. Copy its label, formula, service period, operating window, threshold, measurement point, exclusions and evidence request from the current tender documents. Do not retrofit the historical system to the buyer's vocabulary. Two indicators called availability may count different time, and two four-hour response measures may stop at different events. Similar labels are an invitation to compare definitions, not proof that the values share a meaning.

The fictional Alderpoint Transit Systems Ltd is bidding to monitor passenger-information displays at 48 stations. The buyer asks for evidence of end-to-end availability over a complete year. Its clock covers every minute of the 24-hour service, begins when a synthetic passenger request cannot reach a display and ends after a successful retest. Alderpoint's proposal draft says its existing service achieved 99.982 percent availability in 2025. The source dashboard, however, measured only the central message broker. This is a useful result, but it does not yet answer the buyer's proposition.

The first fields in a service-level evidence record
FieldQuestion to settleAlderpoint entry
Proposed claimWhat exact words will be released?99.982% availability in 2025
Buyer useWhat will the evaluator use it to judge?End-to-end availability for 48 stations
Historical objectWhat actually produced the value?Central message broker on one contract
Immediate issueWhy can the number not travel unchanged?Station displays and telecom links were outside measurement

Turn the label into an event population and a calculation

Identify the unit over which performance is judged. Availability can be time-based, request-based, user-journey based or asset based. Incident performance can count tickets, affected services, customers or distinct outages. Field response can count calls, dispatches, arrivals or restored assets. State the eligible population before looking at the result. Otherwise the analyst can unintentionally choose a denominator that flatters the service.

For a ratio, preserve both parts. If 9,800 of 10,000 eligible requests succeeded, the result is 98 percent only under the recorded success definition. Requests rejected upstream, synthetic tests, retries and malformed calls may or may not belong in the population. For time availability, specify total eligible time, observed downtime and any time removed from both numerator and denominator. Do not mix “downtime removed from the denominator” with “downtime counted as available.” They can produce the same headline while asserting different facts.

Aggregations change meaning. An annual ratio of all events weights busy months more heavily. The average of twelve monthly percentages gives each month equal weight. The average performance across sites gives a quiet station the same weight as a central hub. A percentile answers a distribution question; a mean answers another. Google's SRE guidance distinguishes an indicator from its objective and agreement, and warns that averages can hide the tail. That distinction is essential in a proposal: show which measurement was made before describing whether it was good.

  • Name the unit of analysis and every eligibility condition.
  • Define success and failure in observable terms.
  • Keep numerator, denominator and exclusions separately recoverable.
  • Record aggregation, weighting, segmentation and rounding.
  • Distinguish an observed result from a target and a contractual promise.

Make every start, stop, pause and missing observation visible

A response time has no meaning until its events are named. Receipt by a monitored mailbox, creation in the ticket system, human acknowledgement, valid classification, assignment, remote action, technician arrival, restoration and closure are different timestamps. Choose the pair specified by the historical agreement or measurement rule and preserve any transformation. If a portal imported emails every five minutes, the ticket creation time is not the actual receipt time. If a dispatcher corrected an arrival time after the visit, both values and the reason should remain inspectable.

Define the calendar. State timezone, daylight-saving treatment, service hours, holidays, grace intervals and how events crossing a boundary are handled. ISO 8601 provides an exchange representation for dates and time offsets; it does not prove that two clocks were synchronized or that a business-hours calendar was correct. Near-threshold cases deserve particular scrutiny. A three-minute clock drift can reverse the classification of a 30-minute response.

Now inventory exclusions and pauses. Planned maintenance may be permitted, capped or counted. Time awaiting customer access may pause one contract but not another. A downstream carrier failure may be supplier risk in the new service even if the historical contract excluded it. Preserve the rule that existed during the observation period, when it was approved and how each excluded event was classified. A late spreadsheet filter called “non-service” is not a policy merely because it has been used before.

Clock questions for common service measures
MeasurePossible startPossible endCommon ambiguity
AvailabilityFirst failed observationVerified recoveryComponent or user journey
First responseRequest receivedMeaningful human responseAutomatic acknowledgement
RestorationService impairment beginsService usable againWorkaround versus full repair
On-site attendanceValid call acceptedTechnician arrivesDispatch acceptance versus arrival
ResolutionEligible case opensAgreed outcome deliveredClosure despite reopening

Keep the result connected to the events that produced it

A screenshot proves that a display showed a value at one moment. It rarely proves the underlying population or calculation. Preserve the controlled report, its version, author, approval and period, but also identify the event sources, extraction time, query or calculation logic, field definitions, service calendar and correction log. A reviewer should be able to travel from the proposal sentence to the result, from the result to its numerator and denominator, and from those totals to inspectable records or an authorized assurance statement.

Keep lineage across transformations. Raw monitor events may be deduplicated, joined to an asset register, classified against a maintenance calendar and aggregated by month before a chart is created. Record each step, the responsible system or person, input version and output identifier. W3C PROV offers a useful vocabulary of entities, activities and responsible agents; the point is not to force an ontology into a bid, but to retain who did what to which evidence and when.

Protect confidential operational data. The proposal may cite an approved aggregate, client reference or assurance summary while raw incident descriptions remain in a controlled repository. Record disclosure authority, anonymization, retention and reviewer access. If the buyer requests the evidence, follow the procurement channel and permission boundary. Do not expose customer names, vulnerabilities, user data or infrastructure details merely to make the percentage look inspectable.

Minimum provenance for one released result
LayerRetainWhat it prevents
ClaimExact wording and every proposal locationDifferent meanings around one number
ResultValue, period, unit, aggregation and approvalA dashboard snapshot without context
CalculationInputs, filters, formula, code or workbook versionAn irreproducible percentage
EventsSource identifiers, timestamps and retained correctionsTotals detached from observations
AuthorityService owner, disclosure owner and decision timeAccidental release or unowned interpretation

Review the events most likely to disappear from the percentage

Test the negative space. Compare the expected asset, user or ticket population with what the measurement system saw. A monitor covering 37 of 48 stations has not produced evidence for all 48, even if every observed station performed perfectly. Look for periods when collection stopped, sites entered or left service, identifiers changed, tickets were merged, incidents were recategorized or data was backfilled. Missingness is not automatically failure, but it prevents a confident result until its cause and effect are understood.

Review exclusions as a population, not only one by one. Count their frequency, duration and share of total demand. Separate predeclared contractual exclusions from data-quality removal, test traffic, duplicates, withdrawn requests and discretionary management adjustments. Sample the underlying records. If customer-caused delay forms 30 percent of elapsed time, that fact may matter greatly when the new contract uses a continuous clock. A mathematically valid historical exclusion can still make the result unsuitable for the offered service.

Inspect variation. Report the complete period and material segments before selecting an illustration. Show monthly performance, eligible volume and failures, then analyze sites, channels, severity or load where they affect comparability. Do not choose the best month as a “representative” case. Do not average away a missed contractual threshold. A result can support the statement that annual event-weighted performance was 99.9 percent while failing to support “we met 99.9 percent every month.”

  • Reconcile monitored objects with the authoritative service inventory.
  • Quantify missing intervals and the denominator they could affect.
  • Retain original and corrected event values with reason and approver.
  • Profile exclusions by rule, volume, duration and service segment.
  • Test the full time series before selecting any example.

Decide whether the old service can speak for the new one

Build a side-by-side map of the historic service and the offered service. Compare legal entity, customer type, users, locations, channels, volume and peak load, criticality, technology, service components, operating hours, support tiers, delivery partners, dependencies, metric formula, exclusion regime and observation period. Mark each dimension same, materially similar, different, unknown or not applicable. Do not collapse the map into a generic similarity score. One critical difference can defeat an otherwise close match.

A change does not always invalidate the evidence. A larger service may still use the same proven component, and an improved monitoring design may make a conservative historical claim useful. Explain the logical link. If the offered design adds endpoint monitoring and redundant carrier paths, the broker result can evidence that broker's history while the architecture and test evidence address the new boundary. It cannot be relabelled as the availability of a system that did not exist.

Treat entity and supplier-chain changes explicitly. Performance achieved by an affiliate, incumbent, consortium member or subcontractor belongs to the entity and service that performed it. State that relationship and whether the same people, system, process or supplier will perform the relevant part of the offer. Past performance can inform evaluation, as FAR 15.305 illustrates, but relevance, source, context and trend remain part of the assessment. Ownership of a corporate logo is not operational continuity.

Use decisions for the evidence record
StateMeaningPermitted action
directly_comparableMaterial dimensions and metric matchRelease exact result with source boundary
comparable_with_limitsUseful similarities remain with stated differencesRelease narrowed claim and limitations
recalculation_requiredSource events may support the buyer definitionRecalculate before any numeric claim
source_gapRequired records or definitions are absentHold the claim and seek evidence
exclusion_unresolvedRemoved events may alter the resultResolve or present no result
not_representativeSelected period or segment cannot support normal performanceUse only as a named example
historical_context_onlyTrue result does not evidence the offered propositionDescribe only its historic object
claim_removedNo safe and useful formulation remainsRemove every occurrence

Let the limitation travel with the number

Write the supported proposition before writing persuasive copy. Name the measured object, metric, period and relevant boundary in the sentence or its immediate evidence note. “The central message-distribution component recorded 99.982 percent time availability during calendar 2025 under the incumbent contract's measurement method” is narrower than Alderpoint's draft, but it is defensible. Then state why it is relevant and where it is not equivalent to the proposed end-to-end metric.

Do not bury a material limitation in a distant footnote. A qualifier must be close enough that an evaluator or later reuser cannot detach the percentage from it. Avoid “proven,” “consistently,” “across our services,” “industry-leading” and “SLA achieved” unless the record supports those additional propositions. If every month met a threshold, say so only after testing each month. If the value is an annual aggregate, call it an annual aggregate.

Route the record to the service owner, measurement owner, disclosure authority and bid approver. Contract and commercial reviewers decide whether the offer can promise a target; the historical evidence owner does not. Link every occurrence in the response, executive summary, chart and reference form to the same approved wording. Reopen when the tender definition, source report, calculation, offered scope, architecture, supplier chain, permission or observation period changes.

A true 99.982 percent result becomes a narrower, useful statement

Alderpoint retains the 2025 broker report, monitor configuration, monthly source exports, maintenance calendar, calculation workbook and approval. The value is reproduced as 99.982 percent over the broker's eligible annual minutes after the incumbent contract's permitted maintenance treatment. Monthly results range from 99.941 to 100 percent. The records support the central component and full calendar year. They do not contain station display status, telecom reachability or passenger-request completion.

A second endpoint dataset begins on 1 April 2025 and covers 37 stations. Eleven stations have no comparable endpoint monitor. Carrier incidents are tagged inconsistently, and several missing intervals cannot be classified. The team does not merge that incomplete dataset with the broker figure. It assigns source_gap for end-to-end history and historical_context_only to the broker value in relation to the buyer's exact proposition. The 99.982 percent number remains available for the narrower component claim.

The released answer explains that the broker history evidences stability of one proposed component, identifies the narrower measurement boundary, and presents the new end-to-end monitoring design separately. It does not claim that Alderpoint previously delivered the buyer's 48-station metric. Operations, commercial and contract owners decide the future availability target using the complete offered design, resources, dependencies and price. The evidence record closes only after every copied percentage uses the approved sentence; it reopens if the buyer revises the metric or the solution boundary changes.

  • Supported: the central broker's recorded 2025 result under its historic method.
  • Unsupported: end-to-end availability across all station displays and links.
  • Unresolved: endpoint history for eleven stations and unclassified gaps.
  • Separate decision: the service level that the new offer may commit to.
  • Reopen triggers: buyer metric, architecture, source correction, scope or permission change.

Useful outcomes from service level performance evidence for a bid

  • One claim resolves to one service-level evidence record and a reproducible calculation.
  • The measured service, users, locations, channels, components and operating hours are explicit.
  • Eligible events and time are separated from excluded, missing, duplicated and test records.
  • Clock start, stop, pause, timezone, aggregation and rounding rules are recoverable.
  • Every result retains its observation period, source version and extraction time.
  • Differences between the historical service and the offered service remain visible.
  • An isolated month, best site or central component cannot silently represent normal performance.
  • Approved wording states what the evidence proves and the limitations material to evaluation.
  • Historical evidence never becomes an unowned future service-level commitment.

How to run the work

  1. 01

    Freeze the proposed claim and its evaluation use

    Record the exact sentence, tender question, scoring use, buyer definition, offered service and consequence if the statement is wrong.

  2. 02

    Define the historical service boundary

    Identify the entity, contract, service, customer group, sites, channels, components, hours, suppliers and period that generated the result.

  3. 03

    Reconstruct the indicator

    State the unit of analysis, eligible population, success event, failure event, numerator, denominator, target, aggregation and rounding.

  4. 04

    Rebuild the clock and exclusions

    Trace start and stop events, pause rules, service calendars, timezones, maintenance, dependencies, duplicates, reopenings and missing observations.

  5. 05

    Verify the source trail and calculation

    Preserve raw events or controlled reports, extraction logic, versions, corrections and approvals, then independently reproduce the published value.

  6. 06

    Test representativeness

    Inspect performance by month, site, severity, channel and load so a selected slice or average cannot conceal unstable or absent service.

  7. 07

    Compare history with the offered service

    List every material change in scope, architecture, geography, demand, operating hours, subcontractors, metric and contractual treatment.

  8. 08

    Authorize bounded wording and monitoring

    Choose the supported use, secure disclosure and service-owner approval, link all occurrences, and set expiry and reopening events.

Questions that change the decision

  • What exact proposition will the evaluator infer from the number?
  • Does the source measure a user outcome, an end-to-end service, one component or a proxy?
  • Which requests, tickets, minutes, assets or visits were eligible to enter the calculation?
  • Which event started, paused, restarted and ended the measurement clock?
  • Were exclusions stated before observation or applied later to improve the result?
  • How much of the expected population is missing, unmonitored or manually corrected?
  • Do monthly and segment results support the aggregate or reveal material variation?
  • Is the historic service sufficiently similar to the entity, scope and operating model offered?
  • Which limitations must travel with the result for an evaluator to interpret it correctly?
  • Who may approve disclosure and who may authorize any future commitment?

Where teams lose control

01

A component uptime figure may be presented as end-to-end service availability.

02

A yearly average may conceal failed months or sites with no monitoring.

03

The denominator may omit failed transactions that never reached the monitored component.

04

Pause and exclusion rules may remove the events users experienced as failure.

05

A response claim may measure acknowledgement, assignment, dispatch, arrival or restoration interchangeably.

06

Reopened or duplicated incidents may be removed inconsistently.

07

Timestamps from unsynchronized systems may create false precision near a threshold.

08

Manual corrections may be accepted without reason, authority or retained before-state.

09

A successful pilot or best-performing client may be portrayed as ordinary operations.

10

Historical evidence may be mistaken for approval to promise the same target in the new contract.

Measure the finished job

Measure the completed workflow, including review effort and exceptions. Output volume on its own is not evidence of a better process.

  • service-level claims with complete evidence records
  • calculations independently reproduced from controlled sources
  • eligible population covered by usable measurement records
  • exclusions with a stated rule, reason and authority
  • manual corrections retaining original value and approval
  • claims tested across time and material service segments
  • historical-to-offered service differences resolved or disclosed
  • proposal occurrences linked to the approved wording
  • claims narrowed or removed before release

Common questions

Is a dashboard screenshot enough to support an SLA claim?

Usually not by itself. Preserve the dashboard version and capture time, but also retain the metric definition, population, event sources, filters, exclusions, calculation, observation period and approval needed to reproduce and interpret the result.

Can we quote the best-performing month?

Only as a clearly identified month for a legitimate reason. Do not present it as normal or representative performance. Show the complete period and material variation when the evaluator is judging sustained delivery.

Does 99.9 percent availability have a standard meaning?

No. The value depends on the service boundary, eligible time or requests, success definition, measurement point, maintenance rule, exclusions, aggregation and period. Reconstruct those terms for each source and tender.

How should missing monitoring data be treated?

Identify the missing objects and intervals, investigate why they are absent and test the maximum effect on the result. Do not silently count missing time as available or remove it from the denominator without an authorized rule.

Can customer-caused delays be excluded?

Only according to the applicable historic metric and evidence. Preserve the exact pause or exclusion rule and affected events. Then explain any difference if the buyer's metric assigns that dependency differently.

What if the historic and buyer metrics use different definitions?

Do not relabel the old value. Recalculate from source events if the buyer definition can be applied reliably. Otherwise use a narrower contextual claim, disclose the difference or remove the number.

Does a historical result prove we can promise the same SLA?

No. History is evidence about a defined past service. A future commitment also requires an authorized offered design, demand basis, resources, dependencies, price, measurement regime, remedies and contract approval.

When must the evidence record be reopened?

Reopen it after a change to the tender metric, offered service, historical source, calculation, exclusion decision, monitoring coverage, supplier relationship, disclosure permission or approved wording.

Primary references

Tony Kim

Tony Kim

Founder and CEO

Tony writes about applied AI, dependable product engineering and the systems that turn complex response work into controlled delivery.

Proposal software for source-grounded RFP, RFI, DDQ and questionnaire response work.

Bid, proposal, presales, security and compliance teams. Start with the workflow, constraints and evidence you already have.