8 ms·
| Name | Semi-private eval | Public eval | |--------------------------------------|-------------------|-------------| |
by throwaway71271 2y ago
| Name | Semi-private eval | Public eval |
|--------------------------------------|-------------------|-------------|
| Jeremy Berman | 53.6% | 58.5% |
| Akyürek et al. | 47.5% | 62.8% |
| Ryan Greenblatt | 43% | 42% |
| OpenAI o1-preview (pass@1) | 18% | 21% |
| Anthropic Claude 3.5 Sonnet (pass@1) | 14% | 21% |
| OpenAI GPT-4o (pass@1) | 5% | 9% |
| Google Gemini 1.5 (pass@1) | 4.5% | 8% |
https://arxiv.org/pdf/2412.04604 https://arxiv.org/pdf/2412.04604
- kandesbunzler 2y agowhy is this missing the o1 release / o1 pro models? Would love to know how much better they are
- Freebytes 2y agoThis might be because they are referencing single step, and I do not think o1 is single step.
- aimanbenbaha 2y agoAkyürek et al uses test-time compute.