25 ms·
We have a few sets: 1. Public Train - 1,000 tasks that are public 2. Public Eval - 120 tasks that are public So for those two we don't have protections. 3. S
by gkamradt 1y ago
We have a few sets:
1. Public Train - 1,000 tasks that are public
2. Public Eval - 120 tasks that are public
So for those two we don't have protections.
3. Semi Private Eval - 120 tasks that are exposed to 3rd parties. We sign data agreements where we can, but we understand this is exposed and not 100% secure. It's a risk we are open to in order to keep testing velocity. In theory it is very difficulty to secure this 100%. The cost to create a new semi-private test set is lower than the effort needed to secure it 100%.
4. Private Eval - Only on Kaggle, not exposed to any 3rd parties at all. Very few people have access to this. Our trust vectors are with Kaggle and the internal team only.
- zamadatix 1y agoWhat prevents everything in 4 from becoming a part of 3 the first time the test set is run on a proprietary model, do you require competitors like OpenAI provide models Kaggle can self host for the test?
- gkamradt 1y ago#4 (private test set) doesn't get used for any public model testing. It is only used on the Kaggle leaderboard where no internet access is allowed.
- zamadatix 1y agoSorry, I probably phrased the question poorly. My question is more along the lines of "when you already scored e.g. OpenAI's o3 on ARC AGI 2 how did you guarantee OpenAI can't just look at its server logs to see question set 4"?
- gkamradt 1y agoAh yes, two things 1. We had a no-data retention agreement with them. We were assured by the highest level of their company + security division that the box our test was run on would be wiped after testing 2. We only tested o3 against the semi-private set. We didn't test it with the private eval.
- zamadatix 1y agoMakes sense, particularly part 2 until "the final results" are needed. Thanks for taking the time to answer my question!
- YeGoblynQueenne 1y ago>> We were assured by the highest level of their company + security division that the box our test was run on would be wiped after testing Yuri Geller assured us he was bending the spoons with his mind. Somehow it was only when the Amazing Randi was present that Yuri Geller couldn't bend the spoons with his mind.
- levocardia 1y agoIronically "I have a magic AI test but nobody is allowed to use it" is a lot closer to the Yuri Geller situation. Tests are meant to be taken, that should be clear. And...maybe this does not apply in the academic domain, but to some extent if you cheat on an AI test "you're only cheating yourself."
- Jensson 1y ago> but to some extent if you cheat on an AI test "you're only cheating yourself." You cheat investors.
- anshumankmr 1y agoAnd end users and developers and the general public too... But here is the thing, I feel that even if its rote memorizing why GPT4o couldn't perform just as well on ArcAGI 1 on it or did the "reasoning" help in any way?
- QuadmasterXLII 1y agoAre you aware that OpenAI brazenly lied and went back on its word about its corporate structure, board governance, and for-profit status, and of the opinion that your data sharing agreement is different and less likely to be ignored? Or are you at step zero where you aren’t considering malfeasance as a possibility at all?