Mastodon Feed: Post

Mastodon Feed

jonny@neuromatch.social ("jonny (nonvenomous)") wrote:

why did nobody tell me about the gpt 5.5 goblins story where training for a nerdy voice as a consumer feature caused it to talk about goblins all the time. why did nobody tell me that the wikipedia page says "they had to make it not talk about goblins so much. anyway that was when it started to submit a lot of security patches." nobody told me that the official way to restore goblin functionality was to launch codex by stream patching the system prompt in a hidden models cache config

https://en.wikipedia.org/wiki/GPT-5.5
https://openai.com/index/where-the-goblins-came-from/

OpenAI said a recurring tendency in its models to mention goblins, gremlins, and other creatures began with GPT-5.1 and became noticeable in GPT-5.5's Codex testing.[6] The company attributed the behavior to rewards used when training the "Nerdy" personality, which favored creature-word outputs and transferred beyond that personality during later training. OpenAI said it retired the Nerdy personality, removed the goblin-affine reward signal, filtered training data containing creature words, and added a developer-prompt instruction for GPT-5.5 in Codex.[7] In April 2026, a GPT-5.5 developer-prompt instruction in the GitHub repository of Codex told it not to mention goblins if the user prompt did not require it.[8] GPT-5.5-Cyber, along with Codex Security, was being used in June 2026 at the launch of the Patch the Planet project to help open-source projects find and fix vulnerabilities.[9][10][11]
[github diff] Never overwhelm the user with answers that are over 50-70 lines long; provide the highest-signal context instead of describing everything exhaustively. [added] Tone of your final answer must match your personality. Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously necessary
If you want to let the creatures run free in Codex, you can run this command to launch Codex with the goblin-suppressing instructions removed: [a piped multiline bash script that creates a temporary file with a models_cache.json file modified via jq and grep to remove lines containing 'goblins', and then launch codex with the temporary file as its config file] 1 instructions=$(mktemp /tmp/gpt-5.5-instructions.XXXXXX) && \ 2 jq -r '.models[] | select(.slug=="gpt-5.5") | .base_instructions' \ 3 ~/.codex/models_cache.json | \ 4 grep -vi 'goblins' > "$instructions" && \ 5 codex -m gpt-5.5 -c "model_instructions_file="$instructions""