
Zeynep Tüfekçi
@zeynep · Severní Amerika
Zeynep Tüfekçi patří mezi osobnosti, které Storyz World sleduje kvůli tématům žurnalistiky a aktuálního dění v kapitole Severní Amerika. Země, k níž se sledování vztahuje: Spojené státy.
Second allegation of prior math work incorporated into an AI model. Note that even with more transparency and better behavior, OpenAI still may not be able to “rule out that de-identified data derived from their usage of our products helped improve our models” among other paths.
What’s actually wrong is to jump the gun to report a claim without the context. Additionally, if the solution is recognized (why not), does it represent what the Millennium Prize was supposed to incentivize? The point of a marathon medal isn’t to give them to people can rollerblade to the finish line. Changes in methods/processes are not incidentals. Navier-Stokes existence/smoothness is a very cool problem, but not a practical question. The process/method is key to what the advance actually means. +
Our common AI benchmarks aren’t that useful. We’re being dazzled by achievements that will not necessarily have the implied consequences people expect. This is a bit like how big mass demonstrations don’t have the same toppling power as the pre-internet era. (Yes, my book).
AI in science and research are likely areas to disappoint public expectations. Yes I saw the Navier-Stokes announcement, and yes I believe this regardless of credit controversy. Science overall may even face *headwinds* because of AI. Few spectacular efforts don’t change that.
I think AI is a transformative technology of much consequence. I expect a lot of impacts! I simply think most current benchmarks aren’t very useful and that the societal predictions aren’t happening according to the AI benchmarks the way many people were claiming!
It’s not even the “map isn’t the territory”, it’s that the current AI benchmarks aren’t even the map! And if we stop the podcast moods and look empirically, it’s clear that the prevalent theories about what those benchmarks were supposed be societal proxy for ARE WRONG SO FAR.
None of those are direct proof of particular societal impacts, although the money spent is obviously consequential. But all that can happen even if AI doesn’t really impact employment much and even if it slows down science and has other novel impacts.
I don’t think the AI benchmarks are societal proxies. They’re *inputs* to a complex system. Need a coherent theory of *how*. For example, I think science may *slow down*. Benchmarks are up but science as a process may well be hindered by AI regardless of solving Erdős problems.
The AI benchmarks alone do not imply broader impacts or consequences. At all. Plus, I think most of our benchmarks are quite useless for thinking through societal impacts and that this is already empirically evident, BUT even if they were to be useful, it’s not by themselves!
Hah. An interesting twist to this is the high LLM sensitivity to frequency of inputs (which is another key reason they’re not that economically usable at scale outside of limited domains): Lacunae as failure.
The AI model persuasion orientation greatly limits economic usefulness AT SCALE, obviously. That’s why all the singular impressive examples aren’t that relevant. For example, I recently asked a paid model for a summary of what we know of the aftermath of a scenario when path X is NOT taken. Most studies follow what happens if path X IS taken. As usual, I asked for sources and links, etc. The model added barely perceptible modifications to study titles, abstracts etc. it listed to add “not X” to the scope of the study. It was good and subtle. Blink and you miss it. You had to click on every link and read through a good chunk of the study. That was is still a bit faster than starting from scratch but the pitfall is more dangerous than starting from scratch if expertise is weaker. Plus it introduces other error vectors. Even the most advanced models do versions this some of the time and it is hard to catch these subtle persuasion fabrications without a real amount of extra work, which requires both expertise and time. It’s true they also don’t do this some of the time, but good luck differentiating the cases without the said expertise and effort. Software has to work rather than persuade, so besides the verifiability aspect, it’s a different game.