Went through the comments here and there and one thing to note is that there was a question about who do you think should have won instead. This is a good question because it is possible that all submissions were like this or there were ones that looked just worse. It would be quite useful to know who came close as well in this case. If you knew which submissions were good you could have a process to revoke the prize and give it to someone else in case of fraud or negligence or similar.
Having said that it is also possible that the mistakes and claims were a human error, sure a lot gets ai generated these days but there is a chance in which case the accusation does not look so severe anymore.
Assuming this description is accurate, if they were all like this then none of them should have won. They should have all been disqualified and the organizers should have looked at themselves in a mirror for a long time.
I think the underlying problem here is that no single human brain has enough glycogen in reserve to thoughtfully process all the AI slop. It simply cannot be done by mortals.
I've noticed this over and over again with "professionals actually prefer LLM responses" studies. Typically the human generated responses seem better to me on a quick sample, but if I had to review 50 of them I'd probably start taking lazy shortcuts; using superficial language aptitude or factual comprehensiveness instead of critically reading.
It does seem like the human judges here might have given credit for e.g. a 20pg arXiv paper without actually reading it. I can blame them professionally but emotionally I have nothing but sympathy. I truly hate LLMs.
It was like a bunch of mythos scans through and through which then generated the reports for everyone to implement, not sure if in all orgs though. Mythos was great as it came from the top, i.e. a clear incentive. I think bug reporting otherwise does not reach the engineers unless an incident is raised by the customer care against a responsible team.
It might make even more sense once we get to the point of a wider use of encoding the data into dna. For now we have these few commercial players in the field that cad do it (eg look up dna microfactory for storage archiving), IIRC genomika was saying they can do an MB for a 100-200eur.
I did masters in a similar way, just to get some credentials and fill in the gaps and learn something new. There is an idealistic part to it which is quite romantic as you spend the nights learning and doing the assignments. The structure of such online based learning systems is great for a determined person. However the “other” part of such courses are cheating and ai use. It is depressing to know that the specific credentials prove little because of it. So the only valid signal is: this person did not quit and they know how to write a report, use references. You’d need to test them to fully validate the credential.
The problem is that it is quite difficult to access the published papers is you are not in academia or some company that pays for the access, so AA sort of serves that niche to transfer the knowledge. Training on the other hand is a commercial activity to later rent the model, if this would be purely for open weights I suspect everyone was cool with it.
I do not trust either but you have to at least agree that having some sort of mutually recognised data privacy framework is a good idea because the courts can enforce it then. Saying everything must be from EU is also slightly silly and we should instead have something similar like certification (cyber act ?) to ensure enough competition exists to avoid service degradation. IMO cryptography could be the answer to many privacy related issues for the cross border transfers.
Also these decisions related where the data is stored and which service is used are under control of each commercial org buying them. The risks are assessed at the end of the day and in case of any issues the providers change. Why would a publicly funded org store citizen data in the US is a question regardless of privacy laws though.
Privacy laws are actually one of the very useful things that came out. It is difficult to do the same in the US because of the business lobby. It is crazy that US citizens data can be purchased in the “black” market and the used by the agencies. Leaving tech companies to self regulate is just not viable and it is proven time and time again they cannot do it.
China is building for sure, but Russia? The majority of russia does not have normal roads, full of crappy old trains, and infrastructure which was built in the soviet times. Russia is a joke of a country, oligarchy and kleptocracy rules.
This example is a bit over the top and is more of an edge case, subagents of the same session can use the same VM because what is the point to isolate among them? If at least one subagent is trying to hack you then I would consider the whole session was compromised anyway as you cannot guarantee the agents leaking this among themselves.
Having said that it is also possible that the mistakes and claims were a human error, sure a lot gets ai generated these days but there is a chance in which case the accusation does not look so severe anymore.
reply