Ben and Alex recently published a new report detailing the state of the field for AI for Epistemics and Coordination. With the goal of acting on our reports recommendations, between September 26th and 27th, we (Alex and Ben, with some help from Josh and Harry) ran a weekend hackathon at BlueDot’s San Francisco office focused on AI tools for better decision-making and coordination.
We had 143 applications, accepted 30 people, and had ~20 people participate over the weekend. Our primary goal for the hackathon was for it to be used as a space to develop and refine project ideas. We wanted to see whether the environment would help capable builders develop ideas, get rapid feedback from people with relevant context, and leave with a project worth continuing once the weekend was over.
Our main takeaway: the event format worked decently well for its aims, but we underestimated how much help teams would need choosing worthwhile problems, identifying target users, and coming up with cheap tests to determine whether their proposed tools would actually get used (and be useful).
Why we ran it
We’ve spent the past few months learning and thinking about the field of AI for epistemics and coordination. Here’s why we think this field is important:
AI models are already substantially changing how people and institutions work out what to believe, decide what to do, and coordinate on actions they’ve chosen to take.
As AI models become increasingly capable and integrated throughout society, they will have an effect on both minor day-to-day decisions and far more major decisions.
The full apparatus of AI tools to enable better decisions and coordination will not be built by default. While there are commercial incentives to develop lots of valuable tools (e.g. tools to help with forecasting, scenario-planning, and negotiation), some will no doubt be left behind (or delayed) by solely relying on the market alone.
To understand the current state of the field and what was missing from it, we put together a report. One of our main observations from this was that the space seemed underdeveloped relative to its importance (as briefly described above) and current level of tractability.
Now is a good time to build the field
Here are some reasons we think it’s a great time to start projects in the space:
Today’s AI models are really capable, and they’ll only become more capable as time goes on. In theory this should lead to a variety of projects springing up. For example, many of Forethought’s design sketches look a lot more feasible to execute today than when they were proposed 6 months ago. AIs are rapidly improving at tasks relevant to building these tools, such as software engineering, product design, and so on.
People are becoming increasingly aware that AI models are no longer fumbling chatbots. If demonstrably useful tools are developed with clear users (and pathways to adoption) in mind, they could be widely adopted.
The field has already had success in certain niches, for example see FutureSearch (forecasting), AI-written community notes (collaborative community truth-tracking).
To this end, we figured we would take our own report’s recommendation seriously and run a field-building event as an attempt to bring competent builders and operators into the space.
The rough plan for this experiment was: bring promising people together, give them access to experts in the field with a ton of context and first-hand experience, and see what that enabled.
Introducing the problem and yet-to-be-built solutions.
The event itself
Prior to the event, we came up with three “tracks” to organise projects into:
Institutional capacity: Tools that make organisations more capable and less bureaucratic. See here for concrete examples.
Decision-quality evaluations: More rigorous ways to measure whether AI models actually help people reason better. Here are some specific examples.
Open: Anything else that seemed demonstrably useful and on topic. Some examples of what could be in scope.
What worked
Overall, office hours with expert guests were the highest-value part of the weekend. Mia Taylor (from Forethought), Nathan Young (from Goodheart Labs), and Hans Bøggild and Noah Lloyd (from Alma Grants) joined to facilitate group discussions and have lengthy 1-1 discussions with all participants over the course of the weekend.
A huge thank you to all our expert guests who volunteered so much of their time to this!
Due to their experience in the field, they were able to provide teams with tacit context that couldn’t be gained from reading alone. One session, for example, moved into a concrete discussion on why previous internal attempts to replace document workflows at a specific organisation had struggled and what may be more promising instead.
Nathan Young running a Q&A.
In retrospect we’re glad the event size ended up being focused. We had initially planned around a somewhat larger group. However, in practice we found that the smaller cohort meant we could spend a lot of time talking to individual teams, red-teaming ideas, making introductions, and (sometimes) adapting the schedule as we went along.
For this kind of unusually open-ended event, we now think that something like 15-18 participants may be close to the sweet spot. A larger event would’ve needed a lot more structure and, due to that, would’ve been better suited for more clearly scoped and defined fields.
Team formation required a lot less intervention than we expected. While we considered trying to suggest team matches ourselves based on people’s backgrounds, people just sorted themselves into teams naturally. We think giving everyone some time to organically interact over lunch before the event formally started helped here.
Outcomes from the weekend
The event catalysed (at least - there may be more we are unaware of) two new small groups working on potentially impactful solutions to existing problems in the field. We hope that these groups will persist and go on to refine and scale up their projects further.
The weekend’s most promising projects included:
Two government officials built an AI “tabletop” in which agents representing each stakeholder debate a live policy question, so decision-makers can find blind spots, objections and missing voices in their policy proposals.
Another team worked on a coordination platform for events and research fellowships in which an attendee’s profile, interests and their AI agents link up with everyone else’s, so people find the right collaborators and their agents share project context in one place instead of across scattered chats.
Another team built an eval testing whether AI models can investigate AI incidents as well as human investigators.
Everyone looking locked in.
Deliberation.
What we would improve next time
Teams needed more help with choosing problems before finalising. Several teams locked into a project early and spent the weekend building without first testing whether anyone would use it. Next time we’d encourage teams to spend longer running quick tests before committing, and perhaps giving them a framework at the start (e.g. “We’re helping [specific person or role] make [specific decision] within [their existing workflow]. By the end of the weekend, we’ll test [one key assumption], and the first person we’ll ask for feedback is [named person].”)
We would anchor some teams more strongly on existing proposals. There are plenty of existing project proposals in the field that are waiting for someone to pick them up, so participants without an existing interest area could move more quickly into something that is more immediately impactful.
We want even more high-context people in the room. One of the most promising projects of the weekend (ineligible to win due to one of our judges being on the team) came from people with a pre-existing grounding in AI safety work, which resulted in a more immediately impactful project. Our marketing was mostly promotional messages in Slack and Signal groups. Next time we could target the people we’d most want to see/least likely to be represented at the event (e.g. early-stage founders in related spaces, referrals from talented people we already know).
We should have communicated more clearly throughout the weekend. For example, the start time was unclear, and we underspecified the judging criteria resulting in confusion on day one. There were also some smaller logistics that could have been better managed such as having more vegan food options, and not over-ordering food.
Next steps
We’d love to see more field-building events. We think hackathons work best when focused on specific niches in the field (see here for what we mean by “niches”), or around a particular target user (such as journalists or policymakers, for example).
Workshops with experts could be helpful. Expert workshops would be a good way to get clarity on specific problems and identify who might be able to take ownership of them. These could also be tailored to specific niches.
Soon the field could be ready for a conference. Once the space grows bigger, it would be good to host a conference convening researchers, builders, policymakers and funders to help coordinate projects and ideas across the field.
If any of these sound interesting, BlueDot’s Rapid Grants fund projects like this, and we’d love for you to apply! Submit an application here.
We’d also like to give a big thank you to the judges who helped us make the final decision on the winning projects, as well as to BlueDot Impact for sponsoring this event! The judging panel consisted of: Joshua Landes, Oscar Gilg, Noah Lloyd, Hans Bøggild, as well as ourselves (Ben Norman and Alex Csaky).
If you’re building, funding or thinking about AI tools for better decisions and coordination, we’d love to hear from you.









Very interesting! I would have loved to have been able to attend as this kind of hackathon might have been exactly what I needed to address a few key blockers with the philosophical AI companion that I previously built -- on that note, did anyone build anything similar? I saw that a "tabletop" AI debate simulation was made for policymakers (made something similar as well before, although for ethicists: a meta-ethical roundtable where frontier models would debate and try to converge https://github.com/DaseinB612/meta-ethical-round-table) which is rather adjacent, so would be very interested to see what other project proposals were discussed or what the existing proposals the article mentioned are.