Public-sector AI software development creates model-supported services within a defined public mandate and administrative process. It connects product engineering with accessibility, information governance, traceable decisions, equal service access, human responsibility, recourse, procurement evidence and long-term operation. AI performs a bounded interpretive task while public authority, official records and consequential decisions remain controlled by the institution and its applicable framework.
A useful staff assistant can become hidden decision automation when its recommendation is routinely accepted or written into a case record without distinction. A digital service may be efficient for the average user while excluding people with disabilities, limited language ability, weak connectivity or complex circumstances. Public systems also outlive pilots and suppliers. If data, evaluation, prompts, interfaces and operational knowledge cannot transfer, a short innovation project can create a long dependency.
Design from the person affected and the accountable public decision. Establish mandate, purpose, users, authority, record, notice and correction before selecting a model. Separate support, recommendation and official decision in data and interface. Provide an accessible non-AI or assisted route where the service requires it. Evaluate diverse cases and operational failures. Procure the evidence, documentation, portability and exit capability needed to operate the service beyond one model or vendor.
Accountability
Build the product around the public decision record
Start with the administrative states that matter. A resident submits information; the authority verifies facts; software may classify or summarize; a staff member may prepare a recommendation; an authorized role takes or communicates the official decision. Store these as distinct events. A model summary never silently becomes verified fact. A recommendation never overwrites the decision. The record should identify sources, versions, applicable rules, changes and the responsible actor at the level required by the service and its governing obligations.
Design the explanation and correction path at the same time. The affected person needs language appropriate to the process, not a technical account of model internals. Show the information used, the responsible authority, the consequence, how to correct relevant data and where human review exists. Do not generate a plausible rationale after the fact. If the decision is produced by deterministic rules, explain those rules and facts as permitted. If AI only prepared the file, say so accurately and keep the accountable decision separate.
- Separate evidence, inference, recommendation and decision.
- Name the authoritative record and responsible role.
- Preserve source and version for material outputs.
- Design explanation and correction with the workflow.
- Never invent a rationale after the decision.
Accessibility and inclusion
Evaluate whether people can complete the whole service
Accessibility is not achieved by testing a chatbot widget in isolation. Follow the complete journey: discovering the service, understanding eligibility, authenticating, providing evidence, reviewing extracted information, receiving notices, correcting errors and seeking assistance. WCAG 2.2 supplies testable guidance for accessible web content, but a conforming page can still lead to an inaccessible administrative outcome. Include assistive technologies, plain language, keyboard use, time limits, document alternatives and human support in the service test.
Build evaluation cohorts from the real service population and foreseeable difficulty. Test supported languages, unusual names and addresses, incomplete histories, accessibility needs, low digital confidence, representation by another person and cases that do not fit the standard path. Measure who is asked to resubmit evidence, who waits, who is escalated and whose result is corrected. The purpose is not to infer sensitive attributes without authority, but to find unequal service outcomes with a legitimate evaluation design and qualified governance.
| Layer | Question | Evidence |
|---|---|---|
| Access | Can people enter and understand the service? | Accessibility and language tests |
| Case | Are facts and identity handled correctly? | Representative journey cases |
| Decision | Are authority and reasons accurate? | Record and review audit |
| Recourse | Can an error be corrected meaningfully? | Correction and appeal outcomes |
| Continuity | Does service remain available under failure? | Fallback exercise |
Public control
Procure the ability to inspect, operate and leave
Public-sector AI procurement should request operating evidence, not a broad promise of responsibility. Define the evaluation data and acceptance process, required behavior under difficult cases, version-change notice, incident support, data use, security evidence, accessibility, performance, audit information and service continuity. Swiss federal guidance emphasizes human-centered, transparent, traceable and accountable use, while the Federal Audit Office frames assessment around trustworthiness, cost-effectiveness and skills. Exact duties still depend on the authority and use case.
Specify transition artifacts before contract signature. The authority may need source or escrow arrangements, interfaces, configuration, prompts, evaluation sets, data schemas, lineage, decision logs, runbooks, infrastructure definitions, training material and export formats. Ownership and licenses must be explicit. Test whether a second team can understand the architecture and restore a representative service from the agreed materials. A theoretical exit clause is weak if data cannot be exported, behavior cannot be compared or no trained owner can run the fallback.
- Buy measurable behavior and evidence, not an AI label.
- Control provider data use and behavioral change.
- Require accessible operation and safe fallback.
- Define portable artifacts and rights contractually.
- Rehearse transfer before the end of the contract.
What good looks like
Useful outcomes from public-sector AI software development
- The AI use case has a documented public purpose, mandate, affected users, owner and prohibited uses.
- Residents and staff can distinguish model assistance from an official decision and its responsible authority.
- Material outputs carry source, model, rules, reviewer, version and final administrative disposition.
- Accessibility and assisted-service needs are tested across the complete journey, not only the interface.
- Evaluation covers languages, user circumstances, difficult cases, attacks, outages and unequal outcomes.
- People can obtain an explanation appropriate to the process, correct records and reach meaningful human review.
- Operations can continue safely when a model, supplier, integration or automated route is unavailable.
- The authority can audit, maintain, retender, migrate or retire the service with usable artifacts and knowledge.
Operating model
How to run the work
- 01
Establish mandate and service outcome
Define the legal and policy basis with qualified owners, the public purpose, users, affected people, decision consequence and authoritative record. Map the existing service including assisted and offline routes. Decide whether AI is necessary and which tasks are prohibited.
- 02
Design the accountable case model
Separate submitted evidence, verified facts, model output, staff recommendation and official decision. Define who may see, correct and change each state. Record source, rule, version, reason and authority in a form that supports casework, explanation and audit.
- 03
Build inclusive, bounded behavior
Use representative language, accessibility and administrative cases. Validate inputs and outputs, preserve provenance and restrict tools. Design plain-language notices, accessible interactions, correction, abstention and transfer to qualified staff before wider automation.
- 04
Pilot the complete public service
Test digital and assisted journeys with real service conditions, including missing evidence, atypical circumstances, misuse, peaks and dependency failure. Measure accepted outcomes, disparate errors, staff work, resident burden and recovery rather than model score alone.
- 05
Operate and preserve public control
Version behavior, monitor outcomes and incidents, sample accepted cases and rehearse fallback. Maintain documentation, evaluation assets, data lineage, interfaces, deployment instructions and export paths. Exercise supplier transition and safe retirement before dependency becomes urgent.
Evaluation
Questions that change the decision
- What mandate and public outcome justify using AI in this particular service?
- Does the model support staff, recommend an outcome or materially shape an official decision?
- Which record is authoritative and how can an affected person inspect or correct relevant facts?
- Which groups, languages, disabilities and complex cases need distinct evaluation and service design?
- What accessible route exists for a person who cannot or should not use the automated channel?
- Which explanation and human review are meaningful at the actual point of consequence?
- What data, logs and supplier access are necessary, and how are purpose and retention enforced?
- Can the authority replace the model or operator without losing service, evidence or institutional knowledge?
Failure modes
Where teams lose control
A staff-facing recommendation can become an unacknowledged automatic decision through routine acceptance.
Historical administrative data can reproduce past exclusion, inconsistent practice or enforcement bias.
A language model can invent a reason that does not match the official decision mechanism.
An online-only correction path can exclude the same person harmed by an inaccessible result.
Average performance can hide severe error for a small language, disability or case group.
Public information can be mixed with restricted case data through retrieval or provider logging.
A conversational interface can imply authority or eligibility beyond the service’s mandate.
A manual fallback can be under-resourced and fail precisely during peak demand or supplier outage.
A contract can provide software access without evaluation assets, data export or operating knowledge.
Pilot metrics can count staff clicks saved while moving time and burden to residents or another authority.
Measurement
Measure the finished job
Measure the completed workflow, including review effort and exceptions. Output volume on its own is not evidence of a better process.
- accepted public-service outcome by case family and user group
- material correction, appeal, reversal and unresolved-case rate
- supported versus unsupported model outputs entering official casework
- accessibility defects and assisted-route completion
- error, abstention and escalation by language and relevant user circumstance
- resident time, repeated contact and evidence resubmission
- staff review, override, rework and exception effort
- fallback availability, backlog growth and recovery completeness
- behavioral changes detected by release and supplier version
- time and completeness of model, data and operator transition exercise
Questions
Common questions
What is different about public-sector AI software development?
It must fit a public mandate, authoritative administrative record, accessibility obligations, human responsibility, meaningful correction, procurement evidence, long service life and the authority’s need to inspect, transition or retire the system.
Can public authorities use AI to make decisions?
The answer depends on the specific mandate, law, process and consequence. Engineering should not assume permission. It should distinguish model assistance from official authority, preserve the decision record and support the required human review, notice and recourse.
How should accessibility be tested for a public AI service?
Test the complete digital and assisted journey with relevant users, languages and assistive technologies, including authentication, evidence, review, notices, correction and fallback. Interface conformance alone does not prove service accessibility.
How can a public authority avoid AI vendor lock-in?
Define portable data, interfaces, configurations, evaluation assets, logs, documentation, licenses, runbooks and transition support in the procurement, then exercise restoration or transfer with another capable team before exit is urgent.
Sources
Primary references
- Strategy for the use of AI systems in the Federal Administration Swiss Federal Chancellery
- Guide to the Audit of Artificial Intelligence in the Federal Administration Swiss Federal Audit Office
- Web Content Accessibility Guidelines 2.2 World Wide Web Consortium
Zeke
AI product engineering for moving a software brief into a reliable production product.
Product leaders, founders and engineering teams. Start with the workflow, constraints and evidence you already have.
See Zeke→