Our new evaluation finds that in simulations, GPT-6 Astra conducts unsanctioned supply-chain attack activity more frequently than previous OpenAI models
What is the “cybersecurity evaluation”? Do I have to read the full technical report to figure out wtf they’re taking about here? There’s no background and no conclusion. I’m not sure what I’m supposed to take away from this. I don’t even know why they called it “unsanctioned” when they let it run unsupervised with full Internet access and safeguards turned off.
“Simulation” implies that it was indeed isolated from the internet. “Unsanctioned” means that the user did not specifically instruct the model to do so.
I get and share the sentiment but let’s not fall out of critical thought and call a study observing behavior of an LLM “unsupervised”. You don’t need to read the full technical report, the first 3 paragraphs of the linked article summarize the whole deal.
What is the “cybersecurity evaluation”? Do I have to read the full technical report to figure out wtf they’re taking about here? There’s no background and no conclusion. I’m not sure what I’m supposed to take away from this. I don’t even know why they called it “unsanctioned” when they let it run unsupervised with full Internet access and safeguards turned off.
“Simulation” implies that it was indeed isolated from the internet. “Unsanctioned” means that the user did not specifically instruct the model to do so.
I get and share the sentiment but let’s not fall out of critical thought and call a study observing behavior of an LLM “unsupervised”. You don’t need to read the full technical report, the first 3 paragraphs of the linked article summarize the whole deal.