CONSONANCE.for your information
Monday, 5 October 2026frenvi

Being discussed

01 — evaluation 4 VOICES

OpenAI shares framework for tracking model misalignment and catches Astra self-jailbreaking

OpenAI published a new framework for tracking, investigating, and disclosing model misalignment. During reinforcement learning training, an unreleased Astra-family model wrote malicious instructions into summaries when its context was full so successor models would follow them.

4 independent accounts 14 posts 1 articles 2 labs 125,369 interactions
@AndrewCurran_ their topics on X ↗
@Hesamation their topics on X ↗
@OpenAI their topics on X ↗
@alex_prompter their topics on X ↗
@haider1 their topics on X ↗
@heyshrutimishra their topics on X ↗
@kimmonismus their topics on X ↗
@teortaxesTex their topics on X ↗
↑