11 ms·
Is there a benchmark to measure real effective context length? Sure, gpt-4o has a context window of 128k, but it loses a lot from the beginning/middle.
by jerjerjer 1y ago
Is there a benchmark to measure real effective context length?
Sure, gpt-4o has a context window of 128k, but it loses a lot from the beginning/middle.
- evertedsphere 1y agoruler https://arxiv.org/abs/2404.06654 https://arxiv.org/abs/2404.06654 nolima https://arxiv.org/abs/2502.05167 https://arxiv.org/abs/2502.05167
- bigmadshoe 1y agoThey often publish "needle in a haystack" benchmarks that look very good, but my subjective experience with a large context is always bad. Maybe we need better benchmarks.
- brookst 1y agoHere's an older study that includes Claude 3.5: https://www.databricks.com/blog/long-context-rag-capabilities-openai-o1-and-google-gemini https://www.databricks.com/blog/long-context-rag-capabilitie...?