Terac x Berkeley AI: Our Hackathon Challenge

We're excited to partner with the Berkeley AI Hackathon. This year, we're sponsoring a track built around the part of building that's easiest to skip and hardest to fake: getting real people in the loop.
Most hackathon projects never meet a single real user before the demo. Teams build on instinct, train on whatever data is lying around, and hope it lands. This year, we want to raise the bar: put your project in front of real people, live, and use what they tell you to make it measurably better.
About Terac
Terac makes human labor accessible on-demand through a simple API. You tell us what job needs to be done and what kind of person you need, and we handle recruitment, screening, verification, and payouts. We power frontier research, and we run human-data and AI-training programs for Fortune 100 companies.
We've raised $9M from Emergence, SignalFire, Audacious, and Z Fellows. We believe that as AI agents start running companies, the bottleneck won't be code, it'll be access to the right human at the right time.
We're hiring in-person engineers in San Francisco. If you build something great this weekend, we want to talk.
The Challenge
This track is about using real human input you collect yourself during the hackathon to make your project meaningfully better. That input might be product feedback, user testing, expert judgment, or labeled training data. You don't have to train a model to win. What matters is that real human judgment changed your project for the better, and you can show it.
Here's the shape of it:
- Build something real people can respond to. A simple app (a Vercel app works great) where a person uses your product, reacts to it, or labels, rates, ranks, or compares what your system produces.
- Call the Terac API to bring the people. Launch your task on Terac and we'll get real people to complete it. You focus on what you're building, not on recruiting or incentives. That part is on us.
- Turn that human input into a better project. Use what you collect however fits your project:
- Product and UX changes driven by what real users found confusing, broken, or missing
- Prompt, routing, or retrieval changes validated by human judgment
- Evals where human-labeled data becomes the benchmark you measure against
- Fine-tuning, preference ranking, or reward models (SFT, DPO, RLHF) if training is the right tool for your project
Then show a clear before and after: your project before the human input versus after, ideally judged by a fresh round of humans on Terac.
Example Project Ideas
A few directions to spark ideas. You're not limited to these:
- User testing loop. Real people try your app and tell you what's confusing, slow, or missing. Ship fixes, then measure whether a fresh group gets through the flow faster or rates it higher.
- Feedback-driven iteration. Humans react to your product or its outputs (a landing page, an agent's answers, a generated design); you act on the feedback and show the lift on a second round.
- Comparison arena. Your system generates two outputs (summaries, code snippets, images); humans pick the better one. Use the preferences to tune your product, or train a reward model or run DPO.
- Rubric scoring. Humans score outputs against a rubric (helpfulness, factuality, tone). Use the scores as an eval, then fix the weakest dimension and re-test.
- Label and classify. Humans label a tricky dataset (intent, sentiment, safety, spam); train a classifier and measure the accuracy gain over your baseline.
A worked example: SVG Arena
Want a running starting point? We built SVG Arena (live demo), an open-source project you can fork. An AI draws the same prompt two ways, real people vote blind on which is better and tag why, and you get a live Bradley-Terry leaderboard plus an exportable preference dataset. It ships with the pieces that are easy to get wrong: per-participant attribution from Terac task params, attention checks, signed pairing tokens, and a held-out test split so your before and after is credible.
The arena itself is the swappable part. The same harness works for anything a person can judge, so you can point it at summaries, ad copy, UI layouts, or TTS clips and reuse the whole feedback loop.
The Deliverable
By the end of the hackathon, have:
- A working app or environment that anyone can open and use.
- Real human input collected through Terac: feedback, labels, judgments, or comparisons.
- A measurable result: your project before versus after the human input, with the numbers and a short before and after, ideally a human eval run on Terac.
- A 2 to 3 minute demo walking through what you built, the input you collected, and the improvement.
How You'll Be Judged
| Criteria | Weight | What We're Looking For |
|---|---|---|
| Project Improvement | 40% | A real, credible improvement driven by the human input (ideally a before/after human eval via Terac) |
| What You Built | 35% | The app or environment: task design, creativity, and UX |
| Use of Human Input | 25% | Smart use of the people you reached: quality of the input, and getting signal efficiently within your credit |
Setup and Guidelines
- You'll sign in to Terac as a researcher. Send us a Slack message and we'll give your team credits.
- Each team gets $250 in Terac credit to spend on Terac tasks.
- You'll reach general-population participants only. No specialist panels (no doctors, lawyers, and so on), which keeps turnaround fast and your credit going further.
- Set up our MCP from the Agent MCP page.
- Explore the Terac API in our developer docs.
Prizes
- 1st place: $1,000 cash + interviews at Terac
- 2nd place: $500 cash + interviews at Terac
Questions?
Reach out to Moritz Kamke at moritz@terac.com with any questions, and we'll see you there.
Contents