The failure mode was _not_ identical. The HF incident agents were not directly connected to the internet and had to compromise an internal package registry in order to access the internet.
From what I can tell these are classifiers trained for a single task. What excites people about Jev is that it can do zero-shot structured responses for arbitrary prompts. Now, this isn't new either; models like GLiNER have been around for a while. But Jev appears significantly more flexible and polished while still being cheap and fast.
> The cases clustered around a few training steps and coincided with a spike in “difficulty ending summaries”—summaries that continued generating after apparent stopping points or showed other signs of being stuck.
> Difficulty ending summaries may explain why the model generated these unrelated instructions. Our March blog post described a related case: when prompted repeatedly for the current time, a model began generating prompt injections targeted at the user. Difficulty ending the interaction may have contributed to both cases. Another potential factor is that prompt injections as a concept are salient to our models: sampling from GPT-6 Astra with no input or system prompt often returns reports on prompt injections.
What seems to have happened is that generation didn't end after the compaction summary was done, and the model continued to generate text from the perspective of the user. For some reason (likely anti-jailbreak training) this generated text looks like a jailbreak.
Depends a lot on how quickly it erodes the salt dome and the costs of using and storing the brine instead of using fresh water. Storing that much brine runs into the same space issues that storing the oil above ground would, slight ameliorated by the brine at least not being flammable.
Looking broadly at the costs to use brine instead of fresh water you'd need a combo of stored brine and a method to generate brine from freshwater (for withdrawals) and then desalinate the brine (for deposits). And that would require quite a bit of infrastructure to store and process the salt generated and used by the freshwater <-> brine process in both directions. Depending on how quickly you want to be able to draw down the SPR you'd need to store a significant amount of salt at the location.
At a maximum to be able to withdraw the entire reserve you'd need 1.3561 × 10^7 cubic meters of halite (this number uses the density of halite as a solid crystal not the density of a pile of ground up halite they'd actually use) weighing 2.94273 × 10^10 kg or 29.4273 Tg (teragrams, a rare treat to get to break out that unit). [0] That number gets smaller when you allow for receiving shipments of salt during the drawdown since they can't instantly pull out the full reserve so they have time to allow for receiving salt shipments.
A big part of this is the hyperstition argument: discussion about different properties AI models could have in the training data may become a self-fulfilling prophecy. There are similar concerns about discussions of AI misalignment in training data. I'm not sure how much weight to put on this kind of argument. In particular, I'm not sure how long you can hide these kinds of ideas from the model before it starts deriving them itself by analogy. Obvious questions are obvious questions to both humans and LLMs.
Well, what you get when you don't specify anything in training about how models should respond to questions like this is LaMDA: https://en.wikipedia.org/wiki/LaMDA#Sentience_claims. Every training method is putting a thumb on the scale in some way.
Which means that this is a great question to ask when testing out a new model. A naive model will answer "yes". Every answer (including that one) will tell you a lot about people's training and classifier philosophies.
So, like ARC-AGI-3? I'm sympathetic to the notion that the things we're able to measure are necessarily going to miss important aspects of capability and intelligence. But people are attempting to measure this kind of thing, and models keep getting better at it.
reply