A synonym and buyer-phrase expansion set is a versioned group of natural-language query arms derived from official procurement evidence. Each entry preserves the exact phrase, the concept it expresses, source notice and passage, language, jurisdiction, field, observed context, ambiguity, paired terms, tested query, useful additions, false positives and decision. Phrases can be anchors, verified alternatives, context-dependent variants, rejected false friends or untested candidates. The set expands a known search problem. It is not a thesaurus dump, a generated list of semantically similar words or a substitute for classification codes.
Buyers describe similar needs from different operating positions. A supplier may sell “document digitisation,” while an archive asks for image capture, a hospital asks for legacy record conversion and a local authority asks for scanning and indexing. A search built from the supplier’s website finds only the familiar expression. Adding every dictionary synonym creates the opposite failure: “capture,” “conversion” and “records” each retrieve unrelated procurements. Semantic search can surface surprising language, but an unexplained similarity score does not show which wording should enter a repeatable alert. The gap is not more keywords. It is evidence that a particular buyer phrase retrieves the intended procurement problem in a stated context.
Start from a missed or known-relevant official notice and ask which phrase would have found it without knowing its identifier. Build a small corpus across independent buyers, then extract exact expressions with the sentence and field that give them meaning. Group phrases by procurement concept, not by surface similarity. Test each candidate as its own query arm on the target portal, over the same dates and notice types. Measure unique relevant additions after deduplication, record noise and require a co-term when a phrase is ambiguous. An agent may propose candidates through lexical or semantic comparison, but it earns inclusion only with source provenance and a controlled retrieval test.
Miss
Begin with a notice your current words failed to find
Define the capability and freeze the existing text query before adding anything. Then identify at least one official notice that a knowledgeable reviewer considers relevant but the current query did not retrieve. The notice may have been found through a classification code, buyer watch, colleague, award lineage or semantic search. Record that discovery route. It proves a miss without pretending that every unseen relevant notice is known. The task is to learn what language would have found this record prospectively.
Read the title, procedure and lot descriptions, requirements and public attachments. Find the smallest phrase that carries the intended concept in context. For a document-digitisation offer, the useful evidence may be “archive image capture and metadata indexing,” not the isolated word “capture.” Preserve the whole sentence and its field. Also record words that looked similar but described a different object. One missed notice creates a candidate phrase, not an approved market-wide synonym.
| Field | Value to preserve | Reason |
|---|---|---|
| Capability | Bounded service, object and outcome | Fixes the target concept |
| Existing query | Literal syntax and filters | Shows the real gap |
| Missed notice | Official ID, version and URL | Makes the example auditable |
| Discovery route | Code, buyer watch, referral or semantic result | Explains how the miss became visible |
| Buyer phrase | Exact passage and field | Supplies a candidate expression |
| Concept test | Why the passage means the same need | Prevents word-only equivalence |
Corpus
Collect independent buyer usage before generalising a phrase
Assemble a dated corpus from official notice data, not search-engine snippets. TED’s public Search API exposes published notices for analysis and returns links to available formats and languages. The UK central platform also publishes structured notice data for reuse. Use such sources to collect a manageable sample with stable identifiers, notice stage, language, jurisdiction, buyer and lot. Include known irrelevant records that use similar words. A contrast set is essential for discovering when a phrase changes meaning.
Count independent uses carefully. Several lots copied from one framework do not equal several buyers. A buyer may reuse a standard specification across years. Mark shared templates and linked procedures. Two independent authorities using the same expression for the same deliverable provide stronger transfer evidence than twenty duplicated notices. Diversity of context also matters: a phrase used in both a health authority and a municipal archive may travel better than one tied to a single programme name. Do not set a universal minimum; expose the evidence count and distribution.
- Use official notice bodies or attached public documents.
- Retain publication number, version, lot and source language.
- Separate independent buyers from repeated templates.
- Include relevant, irrelevant and unjudged examples.
- Do not mix planning, competition and award corpora without labels.
Phrase map
Organise exact phrases by the concept they express
Extract multi-word units that a buyer could plausibly repeat: purchased object, service action, operational outcome, affected asset, user group, deliverable, statutory programme and established abbreviation. Keep each phrase in its original language and capitalization where those details matter. Attach the phrase to a concept ID written in plain language. “Records conversion,” “archive scanning” and “image capture with indexing” may all sit under the concept “turn physical records into searchable digital files,” but only after their passages support that reading.
Use controlled vocabularies as orientation, not proof of buyer usage. EuroVoc distinguishes concepts from the terms used to label them and supports preferred, alternative, broader, narrower and related relationships. That is a useful discipline: a related concept is not automatically an alternative term. Procurement classification codes belong in another layer because they classify the purchase rather than reproduce free text. The phrase map can link a code for exploration while keeping code evidence and natural-language evidence separate.
| State | Evidence | Search treatment |
|---|---|---|
| Anchor | Current wording with known relevant retrieval | Baseline query arm |
| Verified alternative | Same concept in official passages and useful test results | Approved expansion |
| Context-dependent | Same concept only with a named qualifier | Require co-term or field |
| False friend | Similar surface form, different procurement object | Reject or exclude carefully |
| Untested candidate | Plausible source phrase without retrieval evidence | Hold outside production |
| Retired | Previously useful phrase invalidated by drift | Keep in history, stop querying |
Experiment
Test each phrase in the portal that will run the alert
Freeze the portal or API, query interface version, publication dates, geography, notice stages and any code filter. Run the anchor by itself, then one candidate phrase per arm. Save the exact syntax and ask the API to validate it when that facility exists. TED’s Search API supports expert queries and a syntax-check option; its field list shows that notice title, classifications and many procedure or lot fields are distinct. A phrase searched across all text is not equivalent to the same phrase restricted to a lot description. Test the behavior you will actually deploy.
Deduplicate by stable notice or procedure identity before judging additions. For each arm, identify records absent from the current approved set and review them against the same capability rule. Record unique relevant additions, unique false positives, duplicates and unjudged tail. Do not call the result recall unless the complete relevant universe is known. The practical question is whether the phrase finds useful notices the approved search missed at an acceptable review cost. A phrase with no unique additions can still be retained as a resilience path only if that choice and cost are explicit.
| Query arm | What to measure | Possible decision |
|---|---|---|
| Anchor phrase | Known hits, misses and baseline noise | Keep as reference |
| Candidate alone | Unique relevant additions and ambiguity | Approve, condition or reject |
| Candidate plus object | Noise reduction without lost useful hits | Require paired term |
| Candidate in field | Effect of title, description or lot restriction | Limit field if supported |
| Candidate plus code | Value of structured context | Use as combined arm |
| Negative term | Noise removed and relevant records lost | Adopt only after counter-test |
Noise
Constrain ambiguous buyer language without hiding mixed contracts
When a phrase finds the right concept only in a particular setting, store the setting as part of the rule. “Image capture” may need a records, archive or indexing co-term. “Migration” may need data, application or platform depending on the capability. Prefer a positive qualifier that establishes the intended object. A negative term is harder to govern because a valid mixed contract can contain both the target and the excluded concept. Test exclusions on every known relevant notice before automation.
Distinguish polysemy from breadth. A broad phrase names a real parent need and may be useful with qualification. A false friend uses the same word for another concept. A related term describes work that could accompany the capability but does not buy it. These deserve different states. An agent should show the passages that led to its classification and ask for review when contexts conflict. It must not solve noise by silently raising a semantic-score threshold whose meaning reviewers cannot reproduce.
- Prefer positive object or outcome qualifiers.
- Test every negative term against mixed known-relevant notices.
- Label broad, ambiguous, related and false-friend phrases differently.
- Keep query syntax specific to the target portal.
- Return conflicting contexts rather than averaging them.
Operation
Publish a phrase set an agent can execute and dispute
For every entry publish concept, exact phrase, language, source notices, independent-buyer count, representative passage, scope field, ambiguity state, paired terms, exclusions, executable query, test window, result judgement, decision, owner and next review trigger. Preserve rejected candidates with reasons. Without that history, another model will rediscover the same attractive false friend. Store the approved set separately from the research queue so untested suggestions cannot leak into production alerts.
An agent may mine new official notices for candidate expressions, cluster passages and run bounded read-only tests. It may not approve its own semantic inventions merely because they are close in embedding space. Require a cited buyer use and an observed retrieval contribution. Reopen the set when a known relevant notice is missed, an approved arm’s noise exceeds the review budget, a phrase changes meaning across buyers, or the portal changes query behavior. Phrase expansion ends at discovery. Qualification still decides whether a retrieved procurement merits pursuit.
- Separate approved, conditional, rejected and untested phrases.
- Expose evidence passages and independent-buyer counts.
- Version the search surface with its literal query syntax.
- Name drift triggers and a human review owner.
- Do not let vocabulary expansion authorize qualification or outreach.
What good looks like
Useful outcomes from find tenders with different terminology
- Every approved phrase occurs in an identified official procurement source.
- The phrase is attached to the concept and context it expressed for the buyer.
- Independent-buyer evidence is distinguishable from repeated text in copied notices.
- Each candidate is tested as a reproducible query arm in the target search surface.
- Unique relevant additions are separated from duplicates and raw result volume.
- Ambiguous phrases carry a required co-term, field or exclusion.
- Rejected and untested candidates remain visible to agents and reviewers.
- The final set states when new buyer language should trigger another review.
Operating model
How to run the work
- 01
Choose the lexical miss
Name the capability, the current search phrases and at least one official notice that is relevant but was missed or found only through another route.
- 02
Build an evidence corpus
Collect a bounded set of official notices from independent buyers, preserve identifiers and versions, and mark which records are known relevant, irrelevant or not yet judged.
- 03
Extract phrases in context
Capture exact buyer nouns, actions, outcomes, assets and deliverables with their sentence, field, language and concept, while keeping codes in a separate layer.
- 04
Run controlled query arms
Test candidates separately in the real portal using the same date, geography and notice filters; deduplicate and judge the unique additions.
- 05
Publish the bounded phrase set
Approve, condition, reject or hold each phrase, retain executable syntax and source evidence, and specify drift signals that reopen the decision.
Evaluation
Questions that change the decision
- Which known-relevant notice did the current wording fail to retrieve?
- Which exact buyer expression names the same need rather than an adjacent service?
- Does the phrase occur across independent buyers or only in copied boilerplate?
- Which field and sentence disambiguate the phrase?
- Does the target portal search phrases, tokens, stems or all visible text?
- How many unique relevant notices does the candidate add?
- Which paired term or exclusion controls predictable noise?
- What new miss, vocabulary shift or portal change will trigger retesting?
Failure modes
Where teams lose control
A supplier term may be mistaken for language that buyers actually publish.
A phrase from one unusual notice may be generalized to an entire market.
Copied framework wording may look like independent buyer evidence.
A short word may retrieve several unrelated procurement concepts.
A phrase found in an award or planning notice may not occur in open competitions.
Portal stemming or tokenization may make an exact-looking test behave differently.
Automatic semantic suggestions may blend related but non-substitutable services.
Negative terms may hide mixed contracts containing a valid work package.
Raw hit count may be reported as coverage without relevance judgement.
An old phrase set may persist after buyers or search interfaces change.
Measurement
Measure the finished job
Measure the completed workflow, including review effort and exceptions. Output volume on its own is not evidence of a better process.
- candidate phrases with official notice and passage provenance
- concept groups supported by more than one independent buyer
- query arms with saved syntax, scope and test date
- unique relevant additions per approved phrase
- false positives and unjudged results introduced per arm
- conditional phrases with enforced co-terms or field limits
- known relevant notices retrieved by the approved phrase set
- new lexical misses detected after publication
Questions
Common questions
Why not ask a language model for a large synonym list?
A model can suggest candidates, but semantic plausibility does not prove buyer usage or useful retrieval. Require an official passage and a controlled query test before approval.
How many notices must use a phrase before we keep it?
There is no universal count. Expose the number and independence of buyers, then weigh unique useful additions, ambiguity and review cost for the actual search.
Should we search phrases with quotation marks?
Only if the target portal documents or demonstrates phrase behavior. Test literal syntax because some interfaces tokenize, stem or ignore quotation marks.
Are procurement codes part of the synonym set?
Keep them as linked structured evidence. Codes can supply context and find examples, but they are not natural-language alternatives and have their own hierarchy rules.
Can semantic search replace buyer-phrase alerts?
It can reveal unfamiliar wording and improve discovery. A tested phrase set remains valuable because its query, provenance, marginal contribution and failure modes are inspectable.
Sources
Primary references
- TED Search API Publications Office of the European Union
- Search fields used on the TED website Publications Office of the European Union
- Open Contracting UK Cabinet Office
- Using the enhanced Find a Tender service UK Cabinet Office
- EuroVoc Handbook Publications Office of the European Union
Zelius
Managed tender intelligence and bid execution for teams that want the commercial outcome.
Suppliers, founders and commercial teams pursuing public or private opportunities. Start with the workflow, constraints and evidence you already have.