← Glossary

benchmark

A standard task set used to compare models. Useful for a rough ranking, weak evidence about your particular codebase.

ai 6 videos 19 mentions

Where it comes up

Training Composer 2

May 14, 2026first said at 3:5112 mentions

“It also uses uh multi-head latent attention for its attention mechanism, which makes it efficient to serve. When deciding on which base model to use, we tested several different open models on several different benchmarks. Some of these loo…”

Cursor 101

Apr 7, 2026first said at 26:343 mentions

“And the chart on the right is one that really matters for day-to-day decisions. This is a performance versus cost on cursor bench, which is cursor's own benchmark that mirrors real-world development tasks. Uh you can see composer two sittin…”

Cursor 201

May 19, 2026first said at 2:451 mention

“Okay. So, one just quick announcement. Uh first first up before we get into it, um you might have seen this yesterday. We we just launched uh our newest version of Composer, our own model here at Cursor. So, this just came out, 2.5. You can…”

Developer Productivity Trends

Jan 27, 2026first said at 16:391 mention

“And here's just like an idea of an adoption plan. Um kind of four steps here. First is just again getting it into the hands of as many developers as you can to establish those benchmarks. Um you know cursor can help with uh enablement if if…”

How Cursor uses Cursor

Apr 9, 2026first said at 33:321 mention

“approaches um bug resolution. Um but Opus 6 and Sonnet are also going to give great explanations and the process. Again, Composer 2 is also a great option. Um it's going to deliver an extremely high level of intelligence and be extremely fa…”

Refactoring Legacy Codebases

May 21, 2026first said at 40:581 mention

“I mean, it's it's constantly changing. Right? The the models, the benchmarks, uh, what's the what's the best performing model? We do publish that information on our social media feeds. Um, one thing you could do is just leverage auto mode i…”