A research team is running a paid study to find professionals who are exceptionally good at designing tasks that expose AI model limitations. The trial asks participants to turn real-world workflows into demanding prompts that push ChatGPT to its breaking point.
The study is part of a broader effort to evaluate where current models break down. By creating challenging prompts, researchers can better understand the weaknesses in today's AI systems.
How the Challenging AI Prompts Study Works
Participants spend about an hour translating a complex workflow from their job or personal life into a demanding prompt. The prompt must require reasoning and real-world lookup — not just simple recall or basic generation.
After crafting the prompt, participants run it in ChatGPT to identify where the model fails. They then refine the prompt until it successfully breaks the system. This iterative process is central to the study's goal of finding model limitations.
Finally, each participant writes a clear grading rubric. This rubric must be detailed enough that a stranger could use it to evaluate any AI's attempt at the task. The rubric ensures consistency in how AI failures are measured.
Why This Trial Matters for AI Evaluation
This initial trial serves a specific purpose: identifying individuals suited for ongoing prompt engineering and evaluation work. The researchers are not just collecting data — they are building a bench of skilled people for future projects.
The approach turns everyday professional knowledge into a testing tool. A lawyer, for example, might design a prompt around contract analysis. A doctor might create one around diagnostic reasoning. Each professional brings unique expertise that generic AI tests cannot replicate.
This matters because standard AI benchmarks often fail to capture real-world complexity. A prompt that works in a lab setting may fall apart when applied to actual professional workflows. This study aims to bridge that gap.
Who Should Consider Participating
The study is open to professionals who understand the limitations of AI in their field. If you have ever used ChatGPT and noticed where it struggles with your specific work, this study wants your insight.
The paid nature of the trial compensates participants for their time and expertise. The one-hour commitment is relatively short, but the impact could be significant for those selected for ongoing work.
For professionals interested in AI evaluation, this study offers a direct path into the field. It is not just about testing AI — it is about shaping how AI is assessed and improved.
Our Take: A Smart Approach to Finding AI Weaknesses
This study represents a practical shift in how we evaluate AI systems. Instead of relying solely on technical experts, it taps into the knowledge of everyday professionals who use AI in real situations.
In our view, this is a smart move. AI models are only as good as their ability to handle real-world tasks. Who better to test that than the people who perform those tasks daily?
The focus on breaking the system is particularly valuable. Most AI testing looks at what models can do. This study asks the opposite question: where do they fail? That perspective is essential for meaningful improvement.
For professionals considering participation, the opportunity extends beyond the immediate payment. Being identified as skilled in prompt engineering could lead to ongoing work in a growing field. As AI becomes more integrated into professional life, the ability to design effective prompts — and identify model weaknesses — will only become more valuable.
To put it plainly: this study is worth watching. It could set a new standard for how AI models are evaluated, and it puts professionals at the center of that process.