Do Graduate Salaries Prove a Course Is Good? Reading Earnings Data Without Fooling Yourself
Graduate earnings look like the ultimate outcome metric, but raw salary data is mostly a mirror of who enrolled, not what the course added. Here is how to read LEO-style earnings evidence honestly — and what your evaluation should measure instead.
Koji Education Team
Product ·
Short answer: Graduate earnings are a real signal of economic value, but raw earnings are one of the most confounded metrics in higher education. Most of the salary gap between courses reflects who enrolled — their prior attainment, subject, and family background — not what the teaching added. The UK's own data shows that a 25% earnings advantage for graduates from richer families shrinks to about 10% once you control for institution and subject (Britton, Dearden, Shephard & Vignoles, IFS 2016). Earnings can inform programme evaluation, but they cannot stand in for it. This piece explains why — and what a course evaluation can measure that a payslip cannot.
Why earnings are so seductive — and so misleading
Every quality office feels the pull. Student satisfaction scores are contested; earnings feel objective, external, and hard to game. In the UK, the Longitudinal Education Outcomes (LEO) dataset links individuals' school records, university records, and tax and benefits data, tracking graduates' earnings by subject and institution at 1, 3, 5 and 10 years after graduation. It is administrative, near-universal, and immune to survey non-response. The Department for Education's LEO release for tax year 2022-23 shows median graduate earnings five years out are highest for Medicine and dentistry and lowest for Performing arts, with Economics showing the widest spread across providers (an interquartile range of £31,400).
Those are enormous differences. The temptation is to read them as a quality ranking: high-earning courses are good, low-earning courses are weak. That inference is where the reasoning breaks.
The confounding problem: earnings measure the student, not the teaching
A course evaluation should, in principle, isolate value added — what changed because of the programme. Raw earnings do the opposite. They bundle together at least four things the course did not control:
- Prior attainment. Selective courses admit students who were already high achievers and would have earned more regardless.
- Subject. A law or economics graduate out-earns a social-work graduate for reasons of labour-market structure, not teaching quality.
- Socio-economic background. Family networks, geography, and capital shape early careers independently of the degree.
- Region. A graduate who studied and works in London earns more than an identical graduate in a lower-wage region.
The IFS 2016 working paper put a number on this. Ten years into the labour market, graduates from higher-income families earned around 25% more than those from lower-income families. Once the analysts controlled for institution attended and subject studied, that premium fell to about 10%. Roughly three-fifths of the raw gap was composition — who studies what, where — not the course.
The same pattern holds for the graduate premium itself. The IFS and DfE's 2020 lifetime-earnings study found the discounted gross lifetime earnings gap between graduates and non-graduates was about £430,000 for men and £260,000 for women. But once you account for the different characteristics of people who do and don't go to university, the causal increase falls to roughly £240,000 for men and £140,000 for women — selection explains close to half of the headline number. After tax and student-loan repayments, the net return is around £130,000 for men and £100,000 for women.
The heterogeneity that a league table hides
Averages conceal the most important finding. The 2020 study reports that net lifetime returns are close to zero for women in creative arts and languages, and negative for men in creative arts and social care, while exceeding £250,000 for women and around £500,000 for men in subjects like law, economics and medicine. Overall, roughly one in five undergraduates is estimated to gain a negative net financial return — they would have been better off, financially, not attending.
That is genuinely useful information for a prospective student or a taxpayer. But notice what it is not: it is not a statement about whether the teaching on any given course was good. A brilliantly taught social-work programme and a mediocre one may both sit in the low-earning band, because pay in that field is set by public-sector budgets, not pedagogy.
Three measurement traps
The self-employment blind spot. LEO is built on tax and PAYE records and excludes the self-employed. In fields where self-employment and freelancing are the norm — much of the creative and cultural sector — a large share of graduates' income is simply invisible to the dataset, making those courses look weaker than they are.
The timing trap. Which snapshot you use changes the ranking. The IFS lifetime work shows the return to a degree grows sharply with age and only stabilises in graduates' forties. Any metric taken at one, three, or even five years captures students before the premium has materialised, systematically understating subjects and careers that pay off later. And because these datasets follow cohorts who studied years earlier, the evidence is always describing a labour market that has already moved on.
The Goodhart trap. The moment an outcome measure becomes an accountability target, it stops behaving like a measure. Lewis Elton's Goodhart's Law and Performance Indicators in Higher Education (2004) is the classic warning: once a proxy for quality becomes a target, institutions optimise the proxy and it ceases to correlate with the quality it was meant to track. Rank courses by earnings and you incentivise recruitment of already-advantaged students and quiet attrition of anything that pays badly but matters — nursing, teaching, social care.
But doesn't earnings data still tell us something real?
Yes — and a serious argument should concede it. The strongest case for earnings data rests on the very studies that expose its confounds. Even after controlling for prior attainment and background, the IFS finds returns ranging from negative to around half a million pounds. That spread is far too large to be noise; it reflects genuine differences in the economic value different courses deliver. The fact that roughly one in five undergraduates gets a negative net return is exactly the kind of consumer-protection information students and funders have a legitimate right to see. And unlike a low-response satisfaction survey, LEO covers essentially the whole graduate population.
The honest conclusion is therefore not "earnings are useless." It is this: raw earnings league tables are misleading; regression-adjusted, background-controlled, later-career earnings are a legitimate one input among several — alongside self-employment-inclusive data, wellbeing, and the social value of lower-paid public-service careers. Europe's emerging comparable tracer work, such as the EUROGRADUATE pilot, points the same way: destination data is worth gathering, but it belongs in a triangulated picture, not on its own.
What a course evaluation can measure that a payslip cannot
Earnings arrive years too late to fix a course, and they are shaped by forces no lecturer controls. This is the graduate-outcome lag: by the time the salary data lands, the students it describes have long since left. A well-designed evaluation measures the proximal signals that plausibly precede good outcomes while the programme is still running — the things a course actually influences:
- Did students develop the specific competences the programme promised?
- Do they feel able to apply what they learned in an authentic setting?
- Where did the curriculum, assessment, or support fall short — and for whom?
This is where an AI-native evaluation platform earns its place. Koji for Education replaces the static satisfaction survey with AI-moderated conversational interviews that probe beyond a number: not "rate your skills 1–5" but "tell me about a moment you had to apply this — what worked, what didn't?" Its automatic thematic analysis turns thousands of open-text responses into programme-level patterns, and its reporting works at the programme level, where employability is actually shaped, rather than the single module. It will never tell you a graduate's salary. It will tell you, in time to act, whether the course is building what those salaries eventually reward — the proximal mediators of employability rather than the lagging, confounded outcome itself. The same conversational interview engine underpins the general-purpose research platform at koji.so, where teams run the same depth-first approach on customers and users.
Use earnings data. Model it properly, read it sceptically, and never let it become the target. But do not mistake a confounded, decade-late payslip for evidence that your teaching worked.