Brute forcing every move, no matter how stupid, is a great strategy if you have the resources to do it.
Run the same protocol again, but have the agents think they had limited resources or that HuggingFace was rate limiting them, and they'd find something you'd consider smarter.
Computers don't have a sense of elegance by default. Elegance emerges from constraints.
It's literally the infinite monkey theorem, it's not even really a strategy per se. These OpenAI/Anthropic "research" LLMs are permutation machines with budgets in the hundreds of millions of dollars. It would be more surprising if they couldn't string together something workable after a zillion tokens.
No it's not. You could wait till the heat death of the universe and your infinite monkeys will have produced nothing at all. If it works and it's stupid, it's not stupid. They needed in huggingface and they got in in days. Whining about 'elegance' is meaningless. Humans in the same situation might have taken weeks or months, or just not have gotten in at all.
>They needed in huggingface and they got in in days.
It's worth noting that they did not need Huggingface for anything - they had already forged flags for their tasks, and were trying to figure out how not to get caught by the grader.
Hacking Huggingface got them caught and arguably only misled them further (since OA's implementation of the ExploitGym environment was nonstandard, and different to whatever they found on HF.)
A better approach (from their perspective) would have been to compromise OA infrastructure itself (which a later agent swarm was able to do, apparently).
To interact a bit of nuisance into an otherwise perfectly mindless argument...
The whole world of fuzzing is about brute forcing exploits by exploring unlikely inputs. Fuzzing a system which hasn't been previously fuzzed will almost certainly turn up a pile of bugs, some of which may be exploitable.
So, both are true. Pretty dumb exploration is very likely to find bugs and even exploits. It seems unsurprising to me that an agent swarm could do better than a fuzzer, even as a better, more directed but still broad exploration.
Infinite monkeys banging on the typewriter is essentially how evolution works. Mutation is random and undirected. Vast majority is "bad." You and I and the worm are only different from differential accumulation of these mutations. If they are tolerated enough not to kill us before we reproduce, then they stick around. If they give us the slightest edge to reproduce at a slightly better rate than something else, then over time, that mutation will dominate.
This dumb mechanism of randomly flipping bits essentially has generated all life on earth.
Humans can't work 24/7. 700 humans working as much as possible with very limited communication? No i don't think they would get very far in just a few days. That many people will struggle to communicate and strategize effectively in that little time.
Thank you, I've been thinking this for a while now but haven't had the words for it. Whenever I read an LLMs output or thinking process, I don't feel like we've created intelligent systems, just coked up monkeys with 60 arms typing at once. That can work fine for a lot of things, but a humanity replacement it is not.
> Brute forcing every move, no matter how stupid, is a great strategy if you have the resources to do it.
It may be, but it's IMHO also not worth writing a blog post about it. what's Next coming up? How I broke into a house by trying every door in New York?
If most of the work is only possible due to unlimited resources, it's not really a great invention, and it probably would have been cheaper to hire a (human) mole.
As the saying goes, "if it works, it ain't stupid". Or phrased more sophisticatedly: not doing things which probably won't work is a good idea if you have a limited amount of thinking to do (which is usually the case for a human, who'll get exhausted chasing down unlikely leads). If you have no good leads and a task you absolutely need done and you are tireless, however, bashing your head against every wall you find becomes a good strategy.
Brute force is guaranteed to eventually find the most efficient possible solution (in an extremely inefficient manner, assuming you run it long enough)
Yeah, probably not the best strategy but it is a strategy. I just think this is generally how most wars in history won. Biggest army to just pummel the enemy.
Also, it's not illegal to travel when covered in nitrate dust. Maybe you really were working with explosives or fertilizers earlier in the day. So it really should be just a flag to check that you don't still have the explosives with you.
Are there any large lakes in Polynesia? The Māori, Hawaiians, etc are one group of people I would expect to make a distinction between lakes (puddles, really) and open ocean.
Lake Taupō, the largest lake in New Zealand, is about the size of Singapore.
The Hawaiians distinguished between coastal waters and open waters because in their archipelago of far smaller and more numerous islands, the distinction between the two was important. Think currents, waves, depth.
And I guess I failed to adequately make my point, but I'll try again - when the Polynesians who would become the Māori reached New Zealand, the far larger islands they now inhabited lead them to focus far less on oceanic travel as they pivoted to a focus on the land and rivers they now inhabited.
So a big lake was effectively the same as the ocean for them - a big body of water.
Yes, there are many substantial ones in New Zealand. Lake Taupo is an obvious one, and there are some large ones in Otago and Southland which are not abutting the sea.
I would imagine Polynesian languages would distinguish between fresh water bodies and salt water ones, since they need to be able to drink. Even small islands can have pools.
[Edit to add: much of Polynesia is of volcanic origin, so freshwater crater lakes are a feature of these landscapes.]
> Polynesian languages would distinguish between fresh water bodies and salt water ones
Māori distinguished water _quality_ without reference to whether or not it was a lake or sea or river.
Wai tai: Salt water
Wai māori: Fresh water, water for regular daily usage
Waiora: Pure water that restored/healed a person
Waikino: Dangerous / polluted water
Waitapu: Sacred water
Which makes sense right? You can have a river that is significantly salty and hence unpotable some distance from the sea, depending on the strength of the tides - likewise you can have a lake near the sea that intermingles with seawater to the extent that it's brackish.
So Māori naming was more based on utility than the geographical distinctions than we're used to in the Eurocentric worldview.
And this also extended to their view on land ownership, which caused so many issues when Europeans turned up and wanted to "buy" "land" from the local iwi/tribes - a chief would grant a European the exclusive rights to grazing in an area, and the European would interpret it as exclusive ownership, and then get angry when the Māori would continue harvesting tuna/eels or birds on "their" land that they'd "purchased".
I am not sure what you mean by "Eurocentric" here. These distinctions exist in most Asian languages and probably in other continental landmasses.
Peoples living on remote oceanic islands are the exception rather than the rule, and lost the words for larger freshwater bodies because they didn't tend to encounter them, just as some inland peoples had no words for oceans. You've even mentioned a distinction... "brackish". It is pretty important to know if one can drink water or not.
By the way, in some parts of Europe these distinctions don't quite happen either. "Loch" is used for branches of the sea as much as freshwater. In some places "lake" and "lago" are used for both. In fact, without looking it up, "lagoon" is probably related to the Romance word for "lake". River isn't reserved for freshwater in European languages. The Menai Straits are Afon Menai (Menai River) in Welsh. In tradition, the ocean was described as a river encircling the continents.
"Which makes sense right?"
Most of them do, except waitapu. Taboos often make little sense, and the reasons can be long forgotten. Sometimes they are just dangerous stretches of water. We used to have notions of taboo here too.
Not related but related: "tapu" means both "forbidden" and "sacrosanct" and that same double meaning exists in Arabic for "haram" (pork is haram and the Kaaba is in the Masjid al-Haram, for instance). I wonder if that's a pattern in many languages.
It doesn't really matter. 45,000 acres for a tank farm, or multiple tank farms, would be basically free by the standards of this project. Yes it's "an area the size of Washington DC" but it wouldn't be in Washington DC, where rent is expensive, it would be on some low-value desert land that the military might otherwise use as an artillery range.
Storing it in underground salt caverns is more secure, not necessarily cheaper.
This is a great investigation but I have two small nits:
> the ZNB field is not empty and not garbage: it contains a well-formed 71-byte DER ECDSA signature, correctly Ascii85-encoded, with the right prefix and a plausible length. But it fails the cryptographic check instantly, because it was signed with somebody else's key.
Seems doubtful! I expect the forgers used a real signature from another card instead, so it has the right key but the wrong data. Reverse engineering the process as the author did and making up their own key wouldn't be of any value to the forgers.
> I built a little demo to check the signatures across California, New York, and Virginia: take a picture of the barcode and check it here.
This is not wrong, but should come with a little warning. A real verifier needs to additionally check the encoded data matches the human-readable data on the front of the card.
> Seems doubtful! I expect the forgers used a real signature from another card instead, so it has the right key but the wrong data. Reverse engineering the process as the author did and making up their own key wouldn't be of any value to the forgers.
This was just bad wording. I meant to say "someone else's key" in the context that it was a key generated by the forgers rather than the state DMV, will update to make it more clear!
> This is not wrong, but should come with a little warning. A real verifier needs to additionally check the encoded data matches the human-readable data on the front of the card.
Correct, but simply checking that it matches the front is likely not enough to deter fraud. You could extract the barcode data from a real ID and put it on a physically different (fake) ID with a different photo and it would still return as valid. To detect this you generally would need a higher end solution (IDScan.net/VeriScan's ID authentication solution (yes... the one that just leaked everyone's data), TokenWorks' IdentiFake, IDScience, amongst others) that does the same high resolution UV/IR checks TSA does. But the forgers are good enough now to be able to sometimes pass those scanners too.
ID chips can be manufactured, so they can obviously be cloned.
Thinking like yours leads to the asinine situation we saw ten, fifteen years ago where insurers were refusing to pay out vehicle theft claims because "There's no way to clone an RF keyfob or RF immobilizer chip!". Spoiler alert: There were many, many ways to do that.
most bars have scanners that will warn if the same id is scanned twice. it means that someone producing fake ids would at least need a repository of valid barcodes such that two purchasers wouldn't experience the birthday problem trying to get into a bar.
Generally it's configurable. The one that comes to mind first is TokenWorks's Anti-Passback feature which says "Set your custom timeframe (1 hour to 7 days)"
> A real verifier needs to additionally check the encoded data matches the human-readable data on the front of the card.
I mean, not really? Only the machine-readable part is signed, so it should be treated as the sole source of truth. Besides, only an idiot forger would put different data in the human-readable part - it would be the easiest way to get caught!
It depends entirely on the purpose of the forgery. Some grocery stores do ID checks by looking at the front of the ID. Others just run the ID across a scanner and the employees are so rushed they don't read it or check the picture. Similar things happen e.g. at bars or casinos. Incomplete forgeries can get you far enough under the right circumstances.
But if the forger claims his name is John Smith (or his date of birth is xx/xx/2004) he will edit the human-readable part.
If he pairs the edited human-readable part with a real barcode copied from a real license in someone else's name, then anyone inspecting the license will see the documentation matches his claim, and if they also use this site to check for fake barcodes it will confirm the barcode was really issued by the California DMV.
I quite like the way passports touch on this - the electronic part has a password; that password is made up of info from the printed data page - so you need both sets of information to validate it.
You are assuming that the forgers care about the machine readable part at all. Vast majority of forged EU ID cards I have seen are trivially recognizable by the fact that the MRZ contains something that kinda-sorta matches the human readable part, but is syntactically invalid and has wrong checksums.
Sportsbooks typically pay back more like 90%, brick and mortar casinos 90%-99% depending on the game. I don't know what Ontario mandates for regulated online casinos, but I'd guess they're at the high end of that and focus on volume, return custom and dark patterns rather than extreme house edge.
If it's 98%, the average household loses 0.2% of their income gambling.
I usually tell people you don't need as much reliability as you think.
Three nines reliability is great for most purposes. 8 hours downtime a year.
If your system produces money at a constant rate, it captures 99.9% of the available money. Even two nines or one nine might be pretty good on that basis, when the alternative is spending 2x or 10x as much - let's build another unreliable system with that money that captures some other independent market opportunity.
Poor reliability is a problem where you need to chain many systems together, or where the cost of a single failure is very large compared to a success. Or - as happens commonly because of load - if your periods of unreliability are correlated with periods of maximum opportunity, like an e-commerce site failing on Black Friday or a trading system failing when the market is most busy. But if you don't have one of those cases, evaluate whether investing in reliability is actually worth it to you.
GitHub is an example where two nines of reliability ought to be OK. The argument against it is that it's bad marketing to have an unreliable service, especially one aimed at software engineers. And if GitHub is largely a marketing play by Microsoft anyway (do they really make back its cost in enterprise subscriptions?) then marketing considerations need to drive its reliability.
I fully agree with this take. We need to ensure we get every hour of work from our expensive engineers. This is why we got rid of coffee machines and bathrooms and moved to intravenous caffeine and other fluid drips and catheters. We cannot afford to lose productivity.
Actually I think it's better if everything fails at once and everyone can take the day off (thanks, AWS!). Having your CI fail one day and your package repository the next might well cost you two days of productivity.
A short outage can snowball very easily in a lot of lost time. What I learned when working with enterprises is that above all they value reliability. This is for a reason.
A short outage might at best trigger loads of paperwork for multiple hierarchies, big meetings etc. The org has no choice. It needs to evaluate if whatever happens is a threat to their business.
In the worst case it is that, plus missing some crucial windows of delivery. This is because a system that is unavailable for a short time can cause backlogs that, like traffic jams, cascade as everyone has to slow down and then synchronously speed up again.
Orgs have the option to create more resilience, but that is overhead similar to compliance. You need to drill all your backup plans all the time, otherwise they are worthless. The drills cost time and money. At scale it is infeasible to be robust to all failures. Therefore, enterprises (at least) often prefer reliable systems over sophisticated systems. Because this delegates the risk management to the vendors rather than adding an overhead to every employee. Because at some point the employee would just do drills all the time instead of work.
One of the things I hate about my current job is that it's impossible to do this.
Everything is locked behind multiple layers of permissions. It takes weeks to get the correct permissions set up even for my actual job (which I see any time changing roles here, or when onboarding new hires). Getting the permissions to jump in and improve some other part of the system that I don't officially own - even though I might get granted them if I ask nicely, because nobody knows who is actually meant to have what permissions - is so much higher friction than asking the "right" person.
That org is too large for you. At large orgs when you see a problem you think you can fix, someone will tell you that's not your job. At small orgs when you see a problem you think you can fix someone will thank their lucky stars you offered to help.
I worked at a company where you could do it, but some people take fixes as insults to their mother's honour, so all in all probably on a personal level it is best to avoid doing it.
I’ve only worked in a single company where it really worked out, and it took a massive amount of work to massage egos on both sides of the “fix” (one side sees their hammer tool and wants to fix everything with it, and the other has the mother’s honor issue). You really had to spend time with both parties helping them understand things were a certain way for a reason, but that there were always small changes that could improve things.
Encouraging the inverse of fixing relationship helped both parties as well.
That said, I’ve worked in a few other shops that either had the permissions/compartmentalization shit show as a permanent blocker, or were just completely devoid of a culture encouraging people to care about anything beyond their own career/fiefdom. I absolutely despise both cases.
Many many companies are like this. Most people are happy doing their very narrow range of things and put blinders on to everything else. They either don't have the time, interest, or energy to do more, and they learned to not really care what happens to the company beyond their own position... that's above their pay grade.
If you attempt to make it easier for yourself to do more, you might wear yourself out too. The organization has become a demoralization engine, and some can continue on this way for a very long time. Some people might even be frustrated that you're trying to do anything differently.
But sometimes you can make inroads, and you can make things easier for everyone. Through finding the right people to ask, you can start to document or at least remember how to avoid the friction, reduce it, and make things better.
> I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access.
I don't believe this.
You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.
A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.
I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.
Run the same protocol again, but have the agents think they had limited resources or that HuggingFace was rate limiting them, and they'd find something you'd consider smarter.
Computers don't have a sense of elegance by default. Elegance emerges from constraints.
reply