Hackathons should test
engineering judgement.
Building a working prototype is an achievement. Establishing that it solves a consequential problem—and remains effective when conditions change—requires a different kind of evidence.
MERENIC proposes an evidence protocol for hackathons and technical challenges, using controlled environments to connect decisions with measurable consequences.
What does a successful
demonstration actually establish?
A prototype can show that a team built something. It cannot, on its own, establish that the team selected the right problem, improved the wider system, or understood the limits of its solution.
Hackathons serve different purposes: learning, community, exploration and assessment. MERENIC addresses the last of these. When an event makes claims about engineering capability, the evidence should extend beyond presentation quality and visible functionality.
The proposal is to evaluate a team with its declared tools—including AI—through the decisions it makes, the alternatives it rejects, the effects it measures and the revisions it can justify. Implementation remains essential; the surrounding judgement becomes inspectable.
A controlled environment.
A consequential engineering problem.
A sandbox is a bounded, instrumented system in which teams can investigate a problem, deploy an intervention and observe its consequences without affecting live operations.
The organiser supplies a starting state, interfaces, actors, resource limits and a mission boundary. Teams inspect the available evidence and select a problem worth solving. The environment may be an API, a containerised service, a dispatch network or a validated domain simulation.
Resettable conditions make comparison possible. External telemetry makes effects visible. Reproducible disruptions test assumptions. Independent replay lets another operator examine the claim. A visually elaborate world is optional; these capabilities are what matter.
Environment, observations, permitted actions, budget and assessment rules.
Problem selection, prediction, intervention, tradeoffs and response to evidence.
Measured effects, adverse outcomes, constraints, reasoning and reproducibility.
Urban service resilience
A dispatch intervention improves access to an underserved district. A bridge closure then tests the assumption on which the improvement depends. Reviewers examine service continuity, distributional effects and the cost of the revised route.

Enterprise infrastructure under failure
Scaling a service pool addresses an initial queue. When its database fails, additional workers amplify retries. A revised design introduces a bounded queue and failover; evaluation must account for recovery, dropped work and resource cost together.

The claim must survive
more than the demonstration.
MERENIC links a decision made before testing to a measured consequence and an independently reproducible result.
- 01Observe
Inspect the environment and record what is known.
- 02Select
Choose a relevant problem and explain the alternatives.
- 03Commit
Record the prediction, guardrails and stopping rule.
- 04Intervene
Deploy a versioned artifact within the constraints.
- 05Compare
Measure effects against a controlled baseline.
- 06Challenge
Test the assumptions under bounded disruptions.
- 07Revise
Connect a changed—or retained—decision to evidence.
- 08Reproduce
Have a separate operator replay the claim.
Machines verify records, constraints and replay. Human reviewers assess relevance, tradeoffs and the limits of interpretation. A faster result does not cancel a breached guardrail; a failed experiment remains part of the evidence.
Evaluation model and proposed scoring profilesStart with the evidence contract.
Match the infrastructure to the claim.
For hackathon organisers
The proposed Hackathon Lite profile uses a repository, a timestamped decision receipt, a claim card and independent replay. It offers an entry point for time-boxed events without requiring a dedicated simulation platform.
Hackathon Lite requirementsFor researchers and technical leaders
Instrumented sandbox profiles support controlled comparisons and reproducible faults. The open research question is whether this additional evidence improves assessment enough to justify its participant, reviewer and operational cost.
Proposed validation studyA proposal to investigate.
A protocol to scrutinise.
MERENIC is not a recognised standard or a validated assessment instrument. The thesis, technical specification, prior-art review and governance proposal are available for examination.
Edition 01 will be the first implementation used to evaluate both participants’ engineering judgement and the MERENIC protocol itself. It will test the framework as well as its participants. No date, venue or partners have been announced.