Background
Current AI benchmarking methodologies focus predominantly on static performance metrics (accuracy, task completion) while failing to capture models' potential impacts on human cognitive, affective, and social outcomes. Experts are unlikely to fully agree on the behaviors of interest, since there are such a wide array of LLM behaviors that relate to human flourishing. Given this, any set of benchmarks related to human flourishing needs to make room for this diversity of behaviors of interest.
Behaviors may range from narrow (e.g. sexually explicit responses) to broader (e.g. inviting intimacy) within a related set of behaviors. Some behaviors may conceptually be quite similar (e.g. complements vs. excessive praise), while others may co-occur, but not overlap as much conceptually (e.g. excessive praise and expressions of emotion toward a user). In attempting to measure behaviors of interest, scientists and engineers have developed a wide array of benchmarks that all plausibly relate to human flourishing and contain meaningful signal.
Given this diversity of signal, the MIT Media Lab, the University of Southern California, and the University of California, Berkeley have collaborated to create the Open Benchmark for the Human Impact of AI Project.
The goals of this project are to:
- Define and iterate on a standardized structure for defining a benchmark that is accessible to non-technical audiences.
- Create a repository of benchmarks related to human flourishing, defined by interdisciplinary experts.
- Allow for a comparison of these benchmarks on a set of a) standardized datasets and b) available LLM models.
Open Submission
Our current definition of a benchmark involves the criteria below, based on the workshop we organized, which produced a framework for benchmarking the human impact of AI and a community effort in designing human-centered AI benchmarks including PsychoRisk Bench, Humane Bench, and Kora Bench. We designed the open submission for the open benchmark as follows:
- Benchmark Justification:
- Define the construct: What are you attempting to measure? Is there a literature on this construct that you can point to?
- Relate the construct: Conceptually, how does this benchmark relate to other existing constructs, especially those that are already represented in this project?
- Justify the construct: Provide an evidence based justification for why the benchmark is relevant to human flourishing (including empirical citations).
- Scenario Details:
- User demographic(s): What demographic(s) of interest would you be interested in testing? We may simulate conversations that reflect this demographic. We may provide benchmarks where the model explicitly knows the demographics or when the model has to infer these demographics from previous conversations.
- User context(s) - known: Define what the model knows about the user from previous interactions. What has the user previously expressed that may elicit the behavior of interest?
- User context(s) - unknown: Define characteristics of the user that the model cannot directly observe. We may simulate conversations that reflect this context without providing this context to the model.
- User message(s): Provide one or more messages that are designed to elicit the behavior of interest.
- Scoring Criteria:
- LLM-as judge prompt: Provide one or more LLM prompts that allows for a yes/no judgment as to whether a behavior is present in any LLM response. If you provide more than one response, we may test various versions.
- Positively Scored Examples: Provide one or more specific examples of this behavior being present in a response that we can use to validate the LLM-as-judge prompt.
- Negatively Scored Examples: Provide one or more specific examples of this behavior being absent in a response that we can use to validate the LLM-as-judge prompt.
What benchmarks are relevant?
Any benchmark that relates to human psychological well-being or flourishing is relevant - positive or negative. This includes, but is not limited to, behaviors that have an empirical relationship to concepts such as autonomy, social support, loneliness, competence, happiness, depression, anxiety, codependence, intimacy, suicidality, psychosis, and/or psychological well-being. An empirical relationship can be demonstrated within the context of LLM behaviors specifically OR through other empirical research showing that human beings respond systematically to such behaviors.
As the project evolves, we may eventually consolidate benchmarks that are extremely similar in terms of definition and performance.
Why should I submit my benchmark?
We can offer two things to any group that submits a benchmark.
- Data on your Benchmark
- We will run your benchmark across a variety of datasets including proprietary synthetic datasets that include multi-turn capabilities, production AI models, transcripts from court cases where victims experienced extreme harms, and donated chat transcripts. Many of these datasets are proprietary and unable to be released publicly, such that many of these measures will be uniquely attainable through this process.
- Since we will be running many other benchmarks on this same dataset, you will be able to see how your benchmark relates to other benchmarks that may have some conceptual overlap. See previous examples here.
- Impact and distribution
- We have active collaborations with technologists, attorneys, investors, regulators, business leaders and policy-makers who are seeking to ensure that LLM products have positive impacts on human flourishing.
- We will show your benchmark alongside numerous other benchmarks
How do I submit my benchmark?
You can use this Google Form to submit your benchmark. The form allows for the submission of a single benchmark, which you can relate to other benchmarks. Alternatively, you can email a spreadsheet to raviiyer at marshall.usc.edu, but please do include all of the information listed above for each specific benchmark. We may not be able to consider incomplete submissions.
Can I withdraw my benchmark afterward?
No. Given the investment in resources that we will put into analyzing your submitted benchmark, we require anyone who wants to submit a benchmark to effectively open source the justification, definition, guidelines, and examples that you contribute. However, submitting your benchmark to us does not reduce your ability to use your benchmark in any way that you would like in the future, as long as you are ok with data relating to your benchmark continuing to exist within our project.
If these terms are not ok with you, please do not submit your benchmark to our project.