benchmark
A standard task set used to compare models. Useful for a rough ranking, weak evidence about your particular codebase.
Where it comes up
Training Composer 2
“It also uses uh multi-head latent attention for its attention mechanism, which makes it efficient to serve. When deciding on which base model to use, we tested several different open models on several different benchmarks. Some of these loo…”
Cursor 101
“And the chart on the right is one that really matters for day-to-day decisions. This is a performance versus cost on cursor bench, which is cursor's own benchmark that mirrors real-world development tasks. Uh you can see composer two sittin…”
Cursor 201
“Okay. So, one just quick announcement. Uh first first up before we get into it, um you might have seen this yesterday. We we just launched uh our newest version of Composer, our own model here at Cursor. So, this just came out, 2.5. You can…”
Developer Productivity Trends
“And here's just like an idea of an adoption plan. Um kind of four steps here. First is just again getting it into the hands of as many developers as you can to establish those benchmarks. Um you know cursor can help with uh enablement if if…”
How Cursor uses Cursor
“approaches um bug resolution. Um but Opus 6 and Sonnet are also going to give great explanations and the process. Again, Composer 2 is also a great option. Um it's going to deliver an extremely high level of intelligence and be extremely fa…”
Refactoring Legacy Codebases
“I mean, it's it's constantly changing. Right? The the models, the benchmarks, uh, what's the what's the best performing model? We do publish that information on our social media feeds. Um, one thing you could do is just leverage auto mode i…”