The track ships 26 specs translated from their Dafny set, reference solutions that all discharge at level 2 (387 obligations, none unproved), a gnatprove harness, and a negative control that rejects wrong bodies, pragma Assume, SPARK_Mode => Off and syntax errors. It also fails any submission where gnatprove emitted no checks at all - a mistake I made once and would rather nobody repeated.
Results, on my own desktop (qwen3-coder 30B, one Radeon R9700, no cloud)
Model alone, their protocol of 5 attempts with prover feedback: 0/26. Every failure the same missing array initialisation, though 22 of 26 wrote a sound loop invariant.
Same model given a short prose description of the idiom, with rounds that improve the prose rather than the code: 25/26. Frozen prose, one attempt each: 19/26.
For scale, nine frontier models average 79.9% on the same tasks in Dafny.
The caveats matter more than the number. One good prompt with no loop got 19 of those 25. Four of the remaining six came from prose written for earlier work rather than anything learned on the night. Four of the held-out tasks were near-relatives of ones the prose had seen, so without them it is 15/22 and 21/22. And one of the two hard tasks (index of maximum) was only cracked after somebody read a known-good answer, so it is not counted.
I have been reading this group’s ancestors since 1999 and using Ada since 1987. SPARK deserved to be in that comparison and nobody had put it there. Corrections to the translations are very welcome - if a spec is unfaithful I would rather hear it here.
I think this Organization you contribute to is a scam…
None of the people in their website link to this “beneficial AI foundation” and some of the pictures are ripped out of other events that have nothing to do with AI.
All the “researchers” who work on the repos don’t link to their academic emails. The few ORCIDs lead to non existing sites and google scholar can be faked. Most use firstname.lastname@gmail (dot) com and their “personal” website is just a random static website generator template with a few notes that could be easily created by an AI agent.
And apparently it’s financed by Vitalik Buterin, the inventor of Etherium Blockchain, but he has never mentioned it anywhere?
There is no website info, legal, privacy policy, cookie policy or the likes.
Like: The “foundation” links to legitimate sources but no legitimate source links to them. Even the (alleged) founder Max Tegmark has never linked this foundation!
Too many red flags!
Edit: After some more research I think this is a SEO-“botnet”. All of the fake accounts link to each other to increase validity and reach. And in between the fake accounts are a few real ones that use the fake engagement to improve their legitimacy.
I may of not graced this forum in a while but some of the older , more bearded and grouchy members have met me at the Adacore conference in 2013 in Paris. I am one person rather than a research org and I have a limited company but I do exist. Good glad to clear that bit up.
After carefuly sizing up groups to scam across the world , I went for the most carreful problem solving provers I could and attempted to fool them on behalf of a guy named in Dune?
I have loved using Ada since being introduced to it in October 1987 in Aberystwyth. Why should I not try an claim a win for Ada?
I do use AI to organise and support things I have not mentioned, it helps because I am autistic, and no one would understand me other wise. I have tried to include as much evidence as I can and I will eventually be releasing the factory structure as opensource and its involvement with gnatprove later, once I have an income.
I’m sorry if I came over the wrong way. I didn’t mean to accuse you of being a bot. But rather warning you that the project you are trying to contribute to is fake and only exists for SEO hacking purposes.