Distilling Human Judgment at Scale: Inside Deep Funding.
$350,000 is being allocated across 98 open source repositories and their 3,677 dependencies through Gitcoin GR24, with $175,000 from the Ethereum Foundation (EF) and $175,000 from Gitcoin. Rather than running another quadratic funding round, this time the round is being decided by an estimate of past contributions from a human jury, and a prediction market to trade those estimates.
The most interesting part of this round is that attempts to manipulate it become profit opportunities for others to capitalise on.
This is the most ambitious application of prediction markets to a real funding decision that’s ever been run on-chain, and the mechanism is worth understanding in detail because if it produces better outcomes, the same template can fork to any DAO, any protocol, and any community facing the same allocation problem.
Let’s discuss.
Why Markets Beat Committees for Funding Decisions
The case starts with the observation that committees don’t bear the cost of being wrong, and markets fix this by making accuracy expensive to abandon. Traders who misprice repos lose money to traders who reprice them correctly, and the system rewards being right rather than being ‘loud.’
Robin Hanson made this case for futarchy decades before the infrastructure existed to run it, and Vitalik continued the thread with ideas like Distilled Human Judgement, where he framed prediction markets as a way to solve the trust problem in scientific, political, and commercial decisions, especially where consensus on who to trust is nonexistent.
Deep Funding is Distilled Human Judgement applied to retroactive public-goods funding, a core engine of the Ethereum ecosystem. The round’s success or failure tests whether markets can scale the judgment of committees and help maintain better decision-making and greater impartiality.
Distilled Human Judgement: The Primitive That Makes This Possible.
The core primitive Deep Funding relies on is distilled human judgment, and understanding it is the difference between seeing the mechanism as “prediction markets for grants” and seeing why this particular shape can scale to problems committees can’t reach.
The setup is straightforward. You have some evaluation mechanism that produces high-quality judgments. In Deep Funding’s case, it’s human jurors at deepfundingjury.com who compare repos pairwise and score them based on contribution to Ethereum.
The problem is that this mechanism is expensive and can only evaluate a limited number of items, which means evaluating all 98 seed nodes pairwise, all 98 originality scores, and weights for all 3,677 dependencies would take an impossible amount of expert time. So instead of evaluating everything, you run a prediction market on the question “if this item is evaluated, what will the result be?” Most items never get evaluated by the underlying mechanism, but the market gives an estimate for all of them.
The items that do get evaluated serve to resolve the market and determine which traders profit and which lose, and because traders don’t know in advance which items will be selected for evaluation, they have an incentive to trade every item according to their genuine beliefs rather than gaming the subset they think will resolve.
The scaling benefit is the obvious one; expert judgment gets stretched across a much larger set than the experts could possibly evaluate directly. But the less obvious benefit is denoising. Evaluation is inherently noisy because it depends on the mood of the evaluator during evaluation, on which specific evaluator is chosen for a given item, and on countless other factors that vary from one assessment to the next. Evaluation can also be biased in ways that aren’t accessible to the broader pool of traders; an individual evaluator might have a personal grudge against a particular project, or be unusually enthusiastic about a niche technical area. Because traders can’t forecast the specific noise that will show up in any specific evaluation, that noise gets averaged out of the market price. And because traders can’t access or forecast biases held by individual evaluators, those biases also get washed out of the aggregate signal. The market price ends up cleaner than any single evaluation would be, even though it’s being calibrated against those individual evaluations.
This is why Deep Funding can credibly produce weights for 3,677 dependencies despite no committee in the world having the bandwidth to evaluate that many items directly. The market does the work of approximation, the jury does the work of resolution on a sample, and the combination produces estimates across the full set that are cleaner than either side could produce alone.
The Methodology in Depth: Three Levels of Allocation
The allocation works through three interconnected markets that together produce a complete weighted dependency graph from Ethereum down to its leaf-node libraries, and the cleanest way to understand how they fit together is to walk through what happens to a single repo.
Each level has its own Pond contest, its own prize pool, and its own market on Seer, which means model builders can compete at any of the three levels independently or build a unified model that produces predictions across all three.
Level 1- is the seed nodes market, which divides the full $350,000 across the 98 repos with weights that must sum to 1 against Ethereum as the target. If Solidity receives a 0.03 weight, that translates to $10,500 of the pool flowing to that node before any further distribution happens.
Level 2 — is the originality market, which assigns each repo a score between 0 and 1 representing how much of the credit belongs to the repo itself versus its dependencies. A score of 0.8 means the repo is mostly original work where dependencies are generic and replaceable, 0.5 means the repo has done substantial work but leans heavily on its dependencies, and 0.2 means the repo is essentially a fork or wrapper where most of the underlying value sits in the libraries it depends on. If Solidity’s originality score lands at 0.6, then $6,300 of its $10,500 stays at the seed node and $4,200 flows down to its dependencies.
You can see this demonstrated in Figure 2.
Level 3 — is the child nodes market, which divides that $4,200 across Solidity’s individual dependencies with weights that again sum to 1, so a dependency that catches a 0.05 weight at level 3 receives $210 from Solidity’s pass-through pool.
The whole structure is one weighted graph where each node’s allocation depends on the level above it, and resolution at every level happens against pairwise comparisons from the human jury at deepfundingjury.com.
Jurors submit judgments like “solidity is twice as important as geth,” the system takes logs of those ratios so they become differences, finds the values that best match the differences by minimizing Huber loss across all pairs, and then exponentiates the result to recover positive weights.
The scoring function used to grade model submissions is the same one that resolves the prediction markets, which means the Pond contest leaderboard and the Seer trading P&L are measuring the same thing through two different feedback channels.
How to Participate
There are two ways to engage with Deep Funding as a model builder, and both are live right now with no waiting period.
The first path is the Pond ML contest, which runs from March 9 through June 29 and offers $10,000 in total prize money split across the three level, $2,500 each at Level 1 and Level 2, $5,000 at Level 3, plus a separate $10,000 writeup prize pool for the best model documentation across all three levels.
The respective due dates for each level are:
Level 1 — June 30th, 23:59 UTC.
Level 2 — June 16th, 23:59 UTC.
Level 3 — June 2nd, 23:59 UTC.
Simply submit a CSV of predicted weights in the format {source, target, weight} where source and target are given to you, and you fill in the weight. The weights are scored against jury data, and the leaderboard updates as new jury data arrives during the round.
Submitting a writeup is mandatory for leaderboard prizes, which is how we ensure the round produces actual research rather than just a stack of opaque CSVs. The writeup needs to be submitted both alongside the model code and at the Gitcoin governance forum thread, and the username on both submissions has to match for prizes to land correctly.
The second path is trading on Seer directly at deep.seer.pm, where any model builder can upload the same CSV used in the Pond contest to take positions on the live markets. Market prices at any moment reflect the aggregate prediction of every model builder currently trading, which means even if you’re not running a model, you can profit by spotting where the consensus has drifted from where you think the jury will actually resolve. The same CSV that enters the Pond contest becomes a Seer trading position, which means one piece of work generates two revenue streams and validates against two different feedback signals simultaneously.
The Incentive Structure: Forgivable Loans and Trading Subsidies
The participation economics are designed so that the downside for any individual model builder is effectively zero.
Trading credits on Seer are structured as 20% pure grant and 80% forgivable loan. The grant portion is yours to keep regardless of how your model performs, and the loan portion only has to be returned out of whatever’s left after trading. If your model loses money, you only return what remains of the 80%, which means your downside from your own pocket is zero. If your model profits, you keep all the profits and return the original 80% loan, which means the upside is uncapped while the floor is protected.
This structure exists specifically to lower the barrier for builders who have never traded prediction markets before, because a model builder who has spent weeks on feature engineering shouldn’t also have to absorb the risk of losing personal capital on their first deployment.
Seer is providing additional trading subsidies to top performers throughout the round, which means model builders who climb the leaderboard receive ongoing capital to keep trading rather than running out of position size after early losses.
Round 1 participants who profited received a 1,500 sUSDS bonus on top of their trading profits, and participants who took losses still received 500 sUSDS for participating, which creates a positive expected value for engaging seriously with the mechanism even before considering the Pond contest prize pool.
Leaderboard position bonuses include:
- 1st 2000 sUSDS
- 2nd 1500s sUSDS
- 3rd 1000 sUSDS
- 4th-10th 500 sUSDS
- Score better than 9999: 200 sUSDS
The structure is deliberately overdetermined toward making participation rational, because the mechanism’s success depends on liquid markets and liquid markets depend on enough participants showing up.
Why Gaming Makes You Money
If a repo decides to inflate its own market price by buying UP tokens on the Level 2 originality market until the score sits at 100% and zero credit flows to dependencies, the human jury later evaluates that repo and lands on a more reasonable 80% originality score, and DOWN tokens that were trading between 0 and 20 cents redeem at 20 cents.
Anyone who bought DOWN against the inflated price profits the difference, and the repo that tried to game its own market loses money on every UP token it bought above the 80-cent mark. The same mechanic applies at Level 1 and Level 3; any attempt to push a market price away from where the jury will eventually resolve creates a profit opportunity for traders who price the repo correctly, and the more aggressively someone tries to manipulate, the more profit accumulates for the participants who spot the manipulation.
Deep Funding turns manipulation into a direct cash transfer from manipulators to the participants who price honestly.
If this works, the implications extend far beyond one Gitcoin round. Every DAO making funding decisions faces the same problem, where evaluation doesn’t scale, committees get captured, and voting optimizes for marketing rather than merit, and the same template that’s allocating $350,000 to Ethereum infrastructure right now can fork cleanly to any allocation problem where expert judgment is the bottleneck.
The contest is open. The markets are live.
What you build on top is up to you.
