Close Menu
FSNN | Free Speech News NetworkFSNN | Free Speech News Network
  • Home
  • News
    • Politics
    • Legal & Courts
    • Tech & Big Tech
    • Campus & Education
    • Media & Culture
    • Global Free Speech
  • Opinions
    • Debates
  • Video/Live
  • Community
  • Freedom Index
  • About
    • Mission
    • Contact
    • Support
Trending

This Sam Altman-Backed Life Insurer Runs Entirely on Bitcoin, and Just Raised $37.5 Million

23 minutes ago

Here’s a Way to Predict When AI Chatbots Will Turn Bad

1 hour ago

Bitcoin (BTC) and ether (ETH) liquidity rebounds a year after $19 billion crypto flash crash

3 hours ago
Facebook X (Twitter) Instagram
Facebook X (Twitter) Discord Telegram
FSNN | Free Speech News NetworkFSNN | Free Speech News Network
Market Data Newsletter
Saturday, October 10
  • Home
  • News
    • Politics
    • Legal & Courts
    • Tech & Big Tech
    • Campus & Education
    • Media & Culture
    • Global Free Speech
  • Opinions
    • Debates
  • Video/Live
  • Community
  • Freedom Index
  • About
    • Mission
    • Contact
    • Support
FSNN | Free Speech News NetworkFSNN | Free Speech News Network
Home»Cryptocurrency & Free Speech Finance»Here’s a Way to Predict When AI Chatbots Will Turn Bad
Cryptocurrency & Free Speech Finance

Here’s a Way to Predict When AI Chatbots Will Turn Bad

News RoomBy News Room1 hour agoNo Comments4 Mins Read1 Views
Share Facebook Twitter Pinterest Copy Link LinkedIn Tumblr Email VKontakte Telegram
Here’s a Way to Predict When AI Chatbots Will Turn Bad
Share
Facebook Twitter Pinterest Email Copy Link

Listen to the article

0:00
0:00

Key Takeaways

Playback Speed

Select a Voice

In brief

  • George Washington University physicists Neil Johnson and Frank Yingjie Huo published a formula that estimates how many good tokens an AI model produces before its first bad one.
  • In the preprint, the formula correctly predicted whether a model would tip immediately or after a delay in 15 of 16 clear-cut cases.
  • The authors propose a parallel monitor that flags when models below a safety threshold.

Physicists at George Washington University have published a formula that estimates how many good answers an AI chatbot will give before it slips into a bad one, and early tests suggest it works.

The study, by Neil Johnson and Frank Yingjie Huo, appeared in the journal Patterns and builds on a preprint, a version posted publicly before formal peer review, first released in February.

Myriad: Which company will IPO next? Click to make your prediction.

Chatbots can answer sensibly for a long stretch and then veer into something harmful, such as bad advice on self-harm or extremist talk, and there has been no simple way to predict when the swerve will happen. The authors argue that existing safety tools often depend on a cloud connection that offline models lack.

Johnson and Huo trace the problem to the attention head, the part of an AI model that decides which earlier words in a conversation matter most when choosing the next one. As a chat grows, the accumulated context pulls that attention toward one cluster of possible answers or another, until it tips.

This is a common pattern exploited by many jailbreakers, and one of the reasons why mayn companies pay attention to system prompts (pieces of text the AI chatbot reads before any query). However, nobody can point out accurately how much effort is required to effectively weaken a model.

Their formula estimates the tipping point, called n, as the number of good tokens—the word fragments a model produces one at a time—that come out before the first bad one. If the conversation already leans toward the bad side, the model tips right away, with an n of zero. If it leans good, the model delivers a run of fine answers and then flips.

In the preprint, the formula picked the right case, immediate or delayed, in 15 of 16 clear-cut tests, or 94%. The researchers ran those tests on six open-weight models, meaning AI systems whose files are public so anyone can download and run them, from OpenAI, EleutherAI, and Meta.

All six sat between 124 million and 410 million parameters, the adjustable numbers inside a model that serve as a rough measure of its size. The published paper reportedly widens the test to seven models of up to 12 billion parameters, which is still small by current standards.

The target is on-device AI, the kind that runs entirely on a phone or laptop with no internet connection, including companion chatbots people talk to like a friend. Google’s experimental AI Edge Gallery app, which Decrypt tested last year, already lets an Android phone run models offline, and nothing typed into it is sent to Google’s servers, and this seems to be a trend that may grow with time as hardware becomes more powerful and smaller AI models become more capable.

A model running offline has no cloud service checking its output, which is the gap the authors want to close. They propose a low-cost monitor that runs in parallel with the model and flags when n* falls below a safety threshold, a bit like a warning light on a car dashboard.

BitcoinBTC · USD

$82,970−2.24%

Oct 3Oct 5Oct 7Oct 8Oct 10

$86.7k$84.7k$82.7k$80.7k

24h HighHigh$82,978

24h LowLow$82,229

VolVol$744.2M

Market projectionsOdds by Myriad

→

They also describe ways to push the tipping point out of reach, such as injecting content into the conversation so n* lands beyond the length of the response. Alignment training, the process of teaching a model to behave, can shift or suppress tipping for specific prompts but cannot remove the underlying mechanism, the authors say.

In April 2025, Decrypt covered an earlier paper from the same pair showing that “please” “and thank you” have a negligible effect on a model’s output, because the model treats polite words as orthogonal, or unrelated in the math, to the substance of a request. That version modeled a single, deliberately simplified attention head.

The preprint’s tests used small models and a 300-token window, or a few short paragraphs of text, and its predictions could be off by one output.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.

Read the full article here

Fact Checker

Verify the accuracy of this article using AI-powered analysis and real-time sources.

Get Your Fact Check Report

Enter your email to receive detailed fact-checking analysis

5 free reports remaining

Continue with Full Access

You've used your 5 free reports. Sign up for unlimited access!

Already have an account? Sign in here

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Telegram Copy Link
News Room
  • Website
  • Facebook
  • X (Twitter)
  • Instagram
  • LinkedIn

The FSNN News Room is the voice of our in-house journalists, editors, and researchers. We deliver timely, unbiased reporting at the crossroads of finance, cryptocurrency, and global politics, providing clear, fact-driven analysis free from agendas.

Related Articles

Cryptocurrency & Free Speech Finance

This Sam Altman-Backed Life Insurer Runs Entirely on Bitcoin, and Just Raised $37.5 Million

23 minutes ago
Cryptocurrency & Free Speech Finance

Bitcoin (BTC) and ether (ETH) liquidity rebounds a year after $19 billion crypto flash crash

3 hours ago
Cryptocurrency & Free Speech Finance

French Committee Backs Stablecoin Swap Tax and Crypto Exit Tax, Then Rejects the Budget

3 hours ago
Media & Culture

SpaceX Found a Way Around the FCC

4 hours ago
Cryptocurrency & Free Speech Finance

Tokenized commodities eye next phase of growth as gold, silver and oil move onchain

4 hours ago
Cryptocurrency & Free Speech Finance

Justin Sun says Tron post-quantum plan on testnet

4 hours ago
Add A Comment
Leave A Reply Cancel Reply

Editors Picks

Here’s a Way to Predict When AI Chatbots Will Turn Bad

1 hour ago

Bitcoin (BTC) and ether (ETH) liquidity rebounds a year after $19 billion crypto flash crash

3 hours ago

French Committee Backs Stablecoin Swap Tax and Crypto Exit Tax, Then Rejects the Budget

3 hours ago

SpaceX Found a Way Around the FCC

4 hours ago
Latest Posts

Tokenized commodities eye next phase of growth as gold, silver and oil move onchain

4 hours ago

Justin Sun says Tron post-quantum plan on testnet

4 hours ago

Mamdani’s $131.5 Million DoorDash Settlement Is Not What He Says It Is

5 hours ago

Subscribe to News

Get the latest news and updates directly to your inbox.

At FSNN – Free Speech News Network, we deliver unfiltered reporting and in-depth analysis on the stories that matter most. From breaking headlines to global perspectives, our mission is to keep you informed, empowered, and connected.

FSNN.net is owned and operated by GlobalBoost Media
, an independent media organization dedicated to advancing transparency, free expression, and factual journalism across the digital landscape.

Facebook X (Twitter) Discord Telegram
Latest News

This Sam Altman-Backed Life Insurer Runs Entirely on Bitcoin, and Just Raised $37.5 Million

23 minutes ago

Here’s a Way to Predict When AI Chatbots Will Turn Bad

1 hour ago

Bitcoin (BTC) and ether (ETH) liquidity rebounds a year after $19 billion crypto flash crash

3 hours ago

Subscribe to Updates

Get the latest news and updates directly to your inbox.

© 2026 GlobalBoost Media. All Rights Reserved.
  • Privacy Policy
  • Terms of Service
  • Our Authors
  • Contact

Type above and press Enter to search. Press Esc to cancel.

🍪

Cookies

We and our selected partners wish to use cookies to collect information about you for functional purposes and statistical marketing. You may not give us your consent for certain purposes by selecting an option and you can withdraw your consent at any time via the cookie icon.

Cookie Preferences

Manage Cookies

Cookies are small text that can be used by websites to make the user experience more efficient. The law states that we may store cookies on your device if they are strictly necessary for the operation of this site. For all other types of cookies, we need your permission. This site uses various types of cookies. Some cookies are placed by third party services that appear on our pages.

Your permission applies to the following domains:

  • https://fsnn.net
Necessary
Necessary cookies help make a website usable by enabling basic functions like page navigation and access to secure areas of the website. The website cannot function properly without these cookies.
Statistic
Statistic cookies help website owners to understand how visitors interact with websites by collecting and reporting information anonymously.
Preferences
Preference cookies enable a website to remember information that changes the way the website behaves or looks, like your preferred language or the region that you are in.
Marketing
Marketing cookies are used to track visitors across websites. The intention is to display ads that are relevant and engaging for the individual user and thereby more valuable for publishers and third party advertisers.