A bid-probability estimate is a dated judgment about the chance of one precisely defined procurement outcome, conditional on the information and pursuit plan available at that time. A defensible estimate begins with a relevant historical base rate, adjusts for verified opportunity-specific evidence, expresses a range and confidence rather than unsupported decimal precision, and records what would change the judgment. It is not a score for team enthusiasm, a promise of revenue or a mathematical fact extracted from the CRM stage.
Bid pipelines are full of exact numbers that nobody can reproduce. One manager enters 70 percent because the buyer likes the team. Another uses 50 percent whenever two bidders are expected. A stage change automatically moves every opportunity to 80 percent, even though mandatory evidence is missing. The forecast rises as bid cost accumulates and rarely falls until the loss arrives. The number then drives revenue, hiring and pursuit priority without showing its definition, evidence or uncertainty. A false point estimate does not make a subjective judgment objective. It hides disagreement, base-rate neglect and optimism behind arithmetic.
Define the event first, then separate outside-view evidence from the particulars of this bid. Use a comparable base rate for the same procedure, stage, market and supplier position. Record opportunity signals as evidence with direction, strength, reliability and date; do not assign universal percentage points to soft impressions. Produce a plausible range, a central planning value and a confidence statement. Test how the estimate moves under material unresolved facts. Update only when evidence or the conditional pursuit plan changes. After outcomes, compare forecast groups with observed results so the organization learns whether its 30s, 50s and 70s mean anything.
Outside view
Define the outcome and earn the right base rate
Write the event so another person could settle it later. “Entity A, bidding alone for Lot 2 of Procurement X, receives the award decision under the current procedure” is different from being shortlisted, admitted to a multi-supplier framework, winning any lot or eventually receiving call-off revenue. Set the forecast date and information stage. A 40 percent judgment before the documents arrive cannot be compared fairly with 40 percent after final pricing. If the company may withdraw or form a consortium, state whether the estimate is conditional on the current pursuit plan.
Begin with the outside view: what happened in a defensibly comparable set before arguing that this opportunity is special. Segment by factors that materially change the process, such as open or invited procedure, new or existing buyer relationship, incumbent or challenger position, geography, sector, contract size, number of lots and forecast stage. Do not create a cohort so narrow that it contains two convenient wins. Report sample size, period, inclusion rule and missing outcomes. Where data is sparse, use a broader range and lower confidence rather than a more elaborate model.
Interpret the base rate carefully. The company’s submitted-bid win rate is conditional on its past qualification decisions. If it now pursues weaker opportunities, that rate will overstate the new cohort. If stopped bids disappear from the data, early-stage conversion may look stronger than it was. If framework admission is counted as a win while revenue requires later call-offs, the event is wrong. The base rate disciplines optimism, but it does not replace opportunity analysis. It provides the starting distribution that the case-specific evidence must move.
| Field | Question | Why it matters |
|---|---|---|
| Outcome | Award, shortlist, framework place or call-off? | Different events have different frequencies |
| Unit | Whole procurement, lot or any acceptable award? | Prevents double counting and ambiguity |
| Bidder | Solo entity, consortium or alternative structure? | Capabilities and eligibility differ |
| Stage | What information was available? | Makes historical comparison fair |
| Date | When was the judgment made? | Creates a stable forecast record |
| Condition | Which pursuit plan and approvals are assumed? | Separates win chance from execution changes |
Inside view
Adjust from evidence without converting impressions into formula points
Create a short evidence ledger across formal eligibility, buyer need, solution fit, delivery credibility, commercial position, competition and procurement process. Each row states the observation, source, date, direction, materiality, reliability and an alternative explanation. “Buyer attended two product sessions” is an observation. It may indicate interest, but also routine market engagement. “The specification matches our product” needs a clause-level fit assessment and comparison with likely alternatives. “Strong relationship” needs recent behavior relevant to this buying decision, not account sentiment.
Do not add fixed percentage points for every positive. Signals are dependent. Buyer access, understanding of need and tailored solution evidence may all arise from one discovery process. Counting each separately triples the same information. Some factors are gates rather than gradients: a mandatory reference gap can make the conditional win chance near zero unless cured, regardless of narrative quality. Other factors change the distribution without a knowable increment. Use the ledger to support judgment, expose disagreement and define scenarios, not to manufacture a scientific-looking total.
Require contrary evidence. For each favorable claim, ask what the strongest credible alternative explanation is and what observation would disconfirm it. Record known weaknesses in the same view as strengths. Unknown competitor count is an uncertainty, not evidence that every bidder is equally likely. Procurement awards are not raffle tickets: eligibility, quality, price, risk and evaluation interact. If the team cannot observe competitors, model plausible competitive scenarios instead of writing 1 divided by an invented field size.
| Signal | Assessment | Control question |
|---|---|---|
| Mandatory fit | Confirmed, curable, uncertain or failed | What exact proof closes the gate? |
| Buyer need | Directly evidenced or inferred | Could the same behavior have another cause? |
| Solution distinction | Relevant and provable or generic | Will the published criteria reward it? |
| Delivery credibility | Capacity and evidence match scope | Which dependency remains outside control? |
| Commercial position | Competitive range with uncertainty | What volume or term moves the price? |
| Competitive field | Observed facts and scenarios | Are signals independent or hearsay? |
| Process | Timetable and procedure stability | Which event could change the field? |
Uncertainty
Express a range, a planning value and the reason for the spread
Use three outputs for different jobs. The plausible range communicates uncertainty. The central value gives finance or portfolio planning one controlled input. The confidence level describes the quality and stability of the evidence, not the chance of winning. A 55 percent central estimate with a 35 to 70 percent range and low confidence says something different from a tightly evidenced 50 to 60 percent range. Never let the central value erase the bounds in executive reporting.
Build the bounds from scenarios that could actually occur. The lower case may combine a capable incumbent, unresolved reference interpretation and no pricing advantage. The central case uses the best current reading. The upper case may assume a favorable clarification, proven differentiation and a smaller competitive field. Do not stack every positive into an implausible best case or every remote threat into the lower bound. Name the two or three uncertainties that create most of the spread and the next evidence that can reduce it.
Official analytical guidance offers a useful discipline even though a bid estimate is a business judgment. The AQuA Book treats uncertainty as inherent in analytical inputs and outputs and distinguishes verifying that work was done correctly from validating fitness for its intended use. Apply both checks: can another reviewer reproduce the evidence and reasoning, and is the estimate suitable for the decision it informs? The 2026 Green Book also warns that some uses of estimated probabilities can introduce spurious accuracy. A range with visible assumptions is often more decision-useful than 63 percent.
Keep bid probability separate from expected contract value, delivery risk and strategic value. Multiplying a large headline contract value by the central probability may be useful for one portfolio view, but only after accounting for lot structure, optional spend, ramp, timing and the actual forecast event. It does not tell the team whether to bid. A lower-probability pursuit can be rational for learning or market entry, while a high-probability contract may still be unattractive. Those are separate authorized judgments.
- Publish lower, central and upper judgments together.
- Describe confidence in the evidence separately from win chance.
- Name the unresolved facts responsible for most of the range.
- Use a central value for planning without presenting it as certainty.
- Keep probability distinct from value, risk and strategic rationale.
Learning
Update on information and test whether past forecasts meant what they said
Estimate independently before group discussion where practical. The bid lead, account owner and delivery or commercial reviewer record a view and evidence without seeing the others’ number. The review then resolves factual disagreement and dependence between signals. Do not average three unsupported guesses into a supposedly defensible forecast. If judgments remain apart, preserve the range and state the contested assumption. Seniority is not new evidence.
Update when something material changes: an addendum alters scope, a reference is accepted, a partner commits, a competitor withdraws through reliable evidence, the buyer publishes clarification, the solution test fails or the price position changes. Record old and new range, trigger, evidence and approver. Do not increase the number because the team submitted, because more money was spent or because the quarter is closing. Freeze the pre-outcome forecast before award information leaks into it, otherwise later calibration becomes meaningless.
Calibrate by grouping forecasts made at the same stage. Among opportunities forecast in the 40 to 50 percent band, did roughly that share win over a meaningful period? Examine reliability separately for sectors, procedure types, incumbent status and forecasters only when samples can support it. Use proper scoring such as the Brier score if the organization has analytical support, but retain simple band charts that leaders can interpret. A favorable overall average can hide systematic overconfidence in large bids and excessive caution in renewals.
Use the result to change the method. Adjust cohort definitions, retire weak signals, require contrary evidence or widen ranges where data is sparse. The Orange Book’s risk guidance treats likelihood as something that may be assessed objectively or subjectively and calls for attention to sensitivity and confidence. That is the right posture here. Subjective judgment is unavoidable; undocumented precision is not. The aim is a forecast whose uncertainty is honest enough to improve decisions and whose history is stable enough to teach the next team.
- Collect independent views before the group anchors on one number.
- Update only for a recorded material change in evidence or plan.
- Freeze the final pre-result forecast before outcome knowledge.
- Compare observed wins with probability bands at like stages.
- Change the model when calibration exposes repeatable bias.
What good looks like
Useful outcomes from estimate bid win probability
- Every probability refers to a named outcome, lot, bidder structure and forecast date.
- Base rates come from a documented comparable cohort rather than the whole CRM.
- Adjustments point to dated evidence and expose unresolved contrary facts.
- Forecasts include a range, central planning value and confidence level.
- Updates have a reason and do not rise automatically with sunk effort.
- Historical calibration reveals optimism, conservatism and weak forecasting segments.
Operating model
How to run the work
- 01
Define the forecast event
Specify procurement, lot, bidding entity or consortium, outcome, stage, date and conditions. Distinguish award from shortlist, framework admission or later call-off revenue.
- 02
Choose a comparable base rate
Select historical bids with similar procedure, stage, market, relationship, supplier position and decision quality, then show sample size and exclusions.
- 03
Build an evidence ledger
Record each favorable, adverse and unresolved signal with source, date, relevance, independence, reliability and expected effect on the outcome.
- 04
Estimate a range and scenarios
Set lower, central and upper judgments, identify the facts driving their spread and use one approved planning value without pretending the range disappeared.
- 05
Update and calibrate
Change the estimate when new evidence arrives, preserve the history, then compare probability bands with actual outcomes after enough decisions accumulate.
Evaluation
Questions that change the decision
- What exact event does “win” mean for this estimate?
- Which completed bids are comparable at the same information stage?
- How much of the apparent win rate reflects selective pursuit rather than market strength?
- Which facts concern eligibility, buyer fit, competition, solution, delivery, price and process?
- Which signals are independent and which merely repeat one source?
- What unknowns create the distance between the lower and upper bounds?
- Which new evidence would materially move or retire the estimate?
- How accurate and well calibrated were prior forecasts made at this stage?
Failure modes
Where teams lose control
A broad historical win rate may ignore procedure, stage, sector or supplier-position differences.
A sponsor’s confidence may be counted several times through dependent signals.
Unknown competitors may be converted into an arbitrary one-divided-by-bidders formula.
Forecasts may rise with effort spent rather than evidence gained.
A central value may be treated as certain revenue while the range is forgotten.
Teams may suppress adverse evidence to protect pipeline optics.
Outcome data may exclude stopped bids and distort the relevant cohort.
Poor calibration may be hidden by changing definitions or forecasts just before the result.
Measurement
Measure the finished job
Measure the completed workflow, including review effort and exceptions. Output volume on its own is not evidence of a better process.
- forecasts with explicit event, stage and date
- base-rate cohorts with sample size and selection rule
- probability records with favorable and adverse evidence
- forecasts expressed with range and confidence
- updates linked to a material evidence change
- observed win frequency by forecast band
- mean forecast error by segment and stage
- forecast stability before outcome notification
Questions
Common questions
Should win probability equal one divided by the number of bidders?
No. That assumes every bidder has an equal chance and that the field is known. Use credible competitor scenarios alongside eligibility, evaluation, solution, price and delivery evidence.
Why use a range if finance needs one number?
Keep the range as the honest decision record and designate one central planning value for the specific finance use. Report both so the convenience input does not become false certainty.
Should probability increase after submission?
Not automatically. Submission may change the conditional event, but effort spent is not evidence of buyer preference. Re-estimate only from material new facts or a changed pursuit plan.
How much history is enough for calibration?
There is no universal count. Use cohorts large enough to avoid conclusions from a few outcomes, show sample size and period, combine bands when sparse and retain lower confidence until more evidence accumulates.
Sources
Primary references
- AQuA Book guidance on analytical quality assurance UK Government Analysis Function
- The Green Book 2026: risk and uncertainty HM Treasury
- The Orange Book: management of risk HM Treasury
Zelius
Managed tender intelligence and bid execution for teams that want the commercial outcome.
Suppliers, founders and commercial teams pursuing public or private opportunities. Start with the workflow, constraints and evidence you already have.