Myself and @ardamgrey spent a little time figuring out how you could set up a simple AI benchmark in a story and came up with this story we’ve put in the library!
You can select which models you want to test in a page and then pick your evaluator LLMs.
Curious to see what results people get with some open source models.
