New

Now in Claude, ChatGPT, Cursor & more with our MCP server

Back to blog
Sector trends10 min read

How to Write a Course Evaluation Software Tender That Does Not Just Re-Buy Your Incumbent

Most European university evaluation tenders are feature checklists copied from the incumbent supplier's brochure. That approach narrows competition, sits awkwardly with EU procurement rules, and reliably buys the last decade of methodology. Here is how to specify outcomes instead.

Koji Education Team

Product ยท

Short answer: Specify what your evaluation system must achieve - evidentiary quality, response representativeness, actionability, lawful data handling - rather than the features your current system happens to have. EU public procurement rules explicitly favour functional and performance-based specifications and warn against requirements that artificially narrow competition. A checklist tender written from an incumbent's feature list is both a weaker procurement and a methodologically conservative one: it guarantees you buy another version of what you already have.

European universities are, for the most part, contracting authorities subject to Directive 2014/24/EU on public procurement. When an institution replaces its course evaluation system, the tender document does something more consequential than selecting a supplier. It freezes a methodology for the length of the contract - typically four to seven years. Whatever the specification demands is what the sector will keep building.

This is why evaluation tenders deserve more methodological attention than they usually receive. They are written by procurement officers with input from IT, and only sometimes with serious involvement from the people who understand validity, non-response and bias.

The checklist problem

The typical failure mode is easy to recognise. Someone exports the configuration options of the current system, converts them into mandatory requirements, and issues them as the technical specification. The resulting document asks for things like "must support paper scanning with OMR", "must support a 5-point Likert scale as the default response format", or "must produce departmental league tables of mean scores".

Two problems follow.

The legal one. The Directive is explicit that specifications should not artificially restrict competition. Recital 74 of Directive 2014/24/EU states that technical specifications "should be drafted in such a way as to avoid artificially narrowing down competition", including through requirements that favour a particular economic operator by mirroring key characteristics of that operator's offering, and it recommends functional and performance-related requirements as the preferred approach. A specification reverse-engineered from one supplier's product is precisely the risk the recital describes.

The methodological one. Every item on that checklist encodes a methodological commitment - and several are commitments the evidence does not support. Mandating a default 5-point scale prejudges a scale design question. Mandating departmental league tables of means requires cross-unit comparability that SET data cannot deliver, and bakes averaging fallacies into the contract. You will have procured, at considerable expense, a system contractually obliged to do the wrong thing.

Specify outcomes, not features

The alternative is to describe what the system must accomplish and let suppliers propose how. Useful outcome-level requirements for course evaluation include the following.

Evidentiary quality. The system must produce evidence sufficient to support programme-level quality judgements and external review under the ESG and ENQA framework. Ask tenderers to demonstrate how their output supports a periodic programme review, not how many chart types they offer.

Diagnostic depth. The system must distinguish what students rated from why. This is the requirement most likely to differentiate genuinely modern platforms from form builders, and it is invisible in a feature checklist because every product technically supports a free-text box. Specify the outcome - open-text feedback must be resolvable into quantified, named themes at programme scale without manual coding - and let suppliers show it.

Representativeness. The system must support monitoring and mitigation of non-response bias, including reporting on who is missing, not merely how many responded. A response rate is not a representativeness measure, and a tender that asks only for the former will get systems optimised for the former.

Actionability. The system must record what was decided in response to feedback and whether it happened - the closing-the-loop requirement that most incumbent systems handle badly or not at all.

Lawful data handling. GDPR compliance, EU-based hosting and transfer arrangements consistent with Schrems II, defined retention and storage limitation, and controls for special category data appearing in free text. Specify data-protection outcomes with evidence requirements, since every supplier will assert compliance.

AI governance. If any tenderer offers AI-assisted analysis or moderation - and in 2026 most will - require disclosure of how it is used, what human oversight applies, and how the supplier addresses the EU AI Act and algorithmic bias. This belongs in the specification, not in a post-award conversation.

Getting the award criteria right

Directive 2014/24/EU makes the most economically advantageous tender the framing for award, and Recital 89 of the Directive frames the choice in terms of the best price-quality ratio - the economically best solution rather than simply the cheapest bid. That flexibility exists to be used, and evaluation software is a good candidate: the cost difference between platforms is trivial next to the cost of running an institution-wide process that produces evidence nobody trusts.

Practical guidance:

  • Weight quality heavily, and define quality criteria in terms of the outcomes above so scoring is defensible.
  • Score demonstrations against your own data or a realistic scenario, not a supplier-controlled demo. Ask each tenderer to analyse the same anonymised set of free-text comments and present the themes. Differences between products become obvious within minutes, and the exercise is objectively comparable across bidders.
  • Record the scoring reasoning. Award decisions must be transparent and justifiable, and quality scores are challenged more often than price scores.
  • Do not over-weight implementation references from identical institutions, which systematically advantages incumbents and large legacy vendors.

Financial thresholds determining which procedure applies are revised periodically by the European Commission, and national transposition varies - confirm current values and your own national rules rather than relying on figures quoted in any article, including this one.

Critics argue: functional specifications are risky and unmanageable

The objection has merit. Outcome-based specifications are harder to evaluate than checklists. A checklist can be scored by an administrator in an afternoon; assessing whether a platform genuinely produces defensible thematic analysis requires methodological judgement that procurement teams may not have on hand. Functional requirements also create more scope for post-award dispute about whether an outcome was met.

There is a second, fairer objection: continuity has real value. Migrating years of historical evaluation data, retraining thousands of staff, and rebuilding integrations with the student information system are genuine costs, and an institution that reasonably concludes its incumbent is good enough is not committing a procurement sin.

The response to both is proportionality rather than purity. Outcome-based specification does not mean specifying nothing concrete - integration, accessibility, language coverage and security requirements should remain hard and testable. It means not converting methodological choices into mandatory technical requirements. And the continuity argument is a reason to weight migration support in the award criteria, not a reason to write a specification only one supplier can meet.

The practical test is simple: read your draft specification and ask whether a supplier with a better method but a different architecture could win. If the answer is no, the tender is not a competition.

Where Koji fits

Koji for Education is built around the outcomes above rather than the feature set of legacy SET tools. AI-moderated conversational interviews address diagnostic depth directly - the moderator probes beyond a rating to reach the reason, which is what makes feedback actionable rather than merely descriptive. Six structured question types (open_ended, scale, single_choice, multiple_choice, ranking, yes_no) cover conventional instrument needs, so migrating an existing question set is straightforward. Automatic thematic analysis and quality scoring deliver the programme-scale open-text requirement without manual coding, and are exactly what a live scoring exercise on your own comment data will expose.

Moderation is standardised and bias-aware, which matters for the AI-governance section of a modern tender: consistency of questioning is a documented property rather than a hope. Formative and mid-cycle collection supports improvement-focused evaluation alongside summative reporting, and closing-the-loop action tracking plus programme- and institution-level reporting meet the actionability and external-review requirements. Data handling is GDPR-compliant and EU-appropriate by design.

We would rather compete on those outcomes than on a checklist, and we would say the same to an institution that ultimately selects a different supplier: a tender written around evidentiary quality produces a better system regardless of who wins it.

Institutions running general user and customer research alongside academic evaluation use the same AI interview engine on the main Koji platform.


Writing or reviewing an evaluation tender? See how Koji for Education maps to outcome-based requirements - and put us in a live scoring exercise against your own feedback data.