Last yr, the staff started experimenting with a tiny mannequin that makes use of solely a single layer of neurons. (Sophisticated LLMs have dozens of layers.) The hope was that within the easiest potential setting they may uncover patterns that designate options. They ran numerous experiments with no success. “We tried a complete bunch of stuff, and nothing was working. It appeared like a bunch of random rubbish,” says Tom Henighan, a member of Anthropic’s technical employees. Then a run dubbed “Johnny”—every experiment was assigned a random title—started associating neural patterns with ideas that appeared in its outputs.
“Chris checked out it, and he was like, ‘Holy crap. This seems to be nice,’” says Henighan, who was surprised as nicely. “I checked out it, and was like, ‘Oh, wow, wait, is that this working?’”
Suddenly the researchers might establish the options a group of neurons have been encoding. They might peer into the black field. Henighan says he recognized the primary 5 options he checked out. One group of neurons signified Russian texts. Another was related to mathematical capabilities within the Python pc language. And so on.
Once they confirmed they may establish options within the tiny mannequin, the researchers set concerning the hairier activity of decoding a full-size LLM within the wild. They used Claude Sonnet, the medium-strength model of Anthropic’s three present fashions. That labored, too. One function that caught out to them was related to the Golden Gate Bridge. They mapped out the set of neurons that, when fired collectively, indicated that Claude was “considering” concerning the large construction that hyperlinks San Francisco to Marin County. What’s extra, when related units of neurons fired, they evoked topics that have been Golden Gate Bridge-adjacent: Alcatraz, California Governor Gavin Newsom, and the Hitchcock film Vertigo, which was set in San Francisco. All informed the staff recognized hundreds of thousands of options—a kind of Rosetta Stone to decode Claude’s neural web. Many of the options have been safety-related, together with “getting shut to somebody for some ulterior motive,” “dialogue of organic warfare,” and “villainous plots to take over the world.”
The Anthropic staff then took the following step, to see if they may use that data to change Claude’s habits. They started manipulating the neural web to increase or diminish sure ideas—a type of AI mind surgical procedure, with the potential to make LLMs safer and increase their energy in chosen areas. “Let’s say we have now this board of options. We activate the mannequin, one in every of them lights up, and we see, ‘Oh, it is interested by the Golden Gate Bridge,’” says Shan Carter, an Anthropic scientist on the staff. “So now, we’re considering, what if we put a little dial on all these? And what if we flip that dial?”
So far, the reply to that query appears to be that it’s essential to flip the dial the correct amount. By suppressing these options, Anthropic says, the mannequin can produce safer pc packages and cut back bias. For occasion, the staff discovered a number of options that represented harmful practices, like unsafe pc code, rip-off emails, and directions for making harmful merchandise.



