Launch is the beginning of evidence
Before launch, teams work with research, prototypes, forecasts, and informed assumptions. Production changes the quality of evidence. Real people arrive with contexts the team did not anticipate. Content meets search behaviour. Integrations encounter edge cases. Operational teams see where customers ask for help. A launch is therefore not the moment a product becomes finished; it is the moment learning becomes richer.
Yet many project measures end exactly there. The date was met, pages were published, features were accepted, and the team moved on. These facts help account for delivery but say little about value. A service can launch on time and make the customer journey worse. Measurement must reconnect the released work to the problem that justified it.
Build a chain from activity to outcome
A useful measurement model separates activity, behaviour, operational result, and business outcome. Publishing clearer eligibility content is an activity. More people choosing the correct application is a behaviour. Fewer avoidable support calls is an operational result. Lower service cost and higher completion may be business outcomes. The chain helps teams see both early signals and ultimate value without pretending one metric explains everything.
Choose measures that can change a decision. Page views may describe demand, but they rarely reveal whether content helped. Completion rate, error frequency, task time, repeat contact, qualified enquiries, successful self-service, retention, or confidence may sit closer to the objective. The right set depends on the service; a small diagnostic tool and a high-consideration purchase should not share a generic dashboard.
Combine behaviour with explanation
Analytics shows what happened at scale. Research helps explain why. A drop-off in a form may indicate confusion, an irrelevant step, a technical failure, a missing document, or a reasonable decision not to continue. Watching sessions, speaking with users, reviewing search terms, and analysing support contacts give meaning to the pattern.
The reverse is also true. A powerful interview is evidence, but it does not establish prevalence. Teams make better choices when quantitative and qualitative sources challenge and complete each other. The aim is not perfect certainty. It is enough confidence to choose the next most valuable action and a way to observe its effect.
Protect the integrity of the measure
Metrics become dangerous when targets encourage teams to optimise the number while weakening the experience. Shorter handling time sounds efficient until advisors rush complex customers. Increased account creation sounds positive until forced registration creates dormant users. Higher click-through may reward exaggerated language that reduces trust later in the journey.
Balance measures and keep context. Pair conversion with cancellation or return rates. Pair self-service with successful completion and accessibility feedback. Segment results where the aggregate hides unequal outcomes. Document changes to tracking and definitions so trends remain interpretable. Measurement should make the organisation more honest, not merely more numerical.
Create a rhythm for response
A dashboard without a decision rhythm becomes decoration. Assign ownership, review a focused set regularly, and agree what will trigger investigation. Product teams should be able to connect a change in evidence to a prioritised experiment, content revision, technical fix, or research question. Larger strategic outcomes may move quarterly; operational and experience signals may need weekly attention.
The Government Service Standard asks teams to define success and publish performance data because a live service should remain accountable to the people it serves. Commercial organisations may not publish every measure, but the discipline still applies. The most valuable post-launch question is not “did it work?” as a final verdict. It is “what is the experience teaching us now, and what will we change because of it?”
Establish the baseline before change
Teams frequently define measures during launch planning, after the old experience has disappeared or tracking has changed. That makes improvement difficult to demonstrate. Capture a baseline while the current journey still exists. Record the metric definition, time period, segments, known data limitations, and external conditions that could affect comparison. Numbers without this context can create more confidence than they deserve.
A baseline can include operational and qualitative evidence, not only analytics. Support reasons, staff workarounds, customer quotes, usability observations, accessibility defects, processing time, and manual rework create a fuller account of the starting point. They also help explain why a stable top-line metric may hide a meaningful improvement elsewhere.
Do not wait for perfect instrumentation before acting. Identify the minimum evidence needed to judge the first release and improve collection over time. Be explicit about gaps. Measurement maturity comes from repeated use: the team notices which signals clarify decisions, removes vanity metrics, and strengthens weak definitions. The baseline is valuable not because it freezes reality, but because it gives future claims a credible point of comparison.