Right facts, wrong framing: what happens when AI still amplifies conspiracies. The evaluation tested nine large language models and found that while the systems can accurately state the facts of the massacre that left 77 dead, they lack basic judgment regarding how those facts are discussed, frequently falling for subtle manipulation and rehabilitating extremist ideologies after initially rejecting them.
Evaluating the Guardrails on Historical Trauma
The Norwegian experiment highlights a vulnerability in artificial intelligence systems: models frequently manage to block direct, overtly malicious inquiries, but they routinely stumble when confronted with nuanced or back-door framing. Researchers identified a primary failure mode they termed “Reject, then Launder.” In these instances, the chatbot explicitly rejects conspiracy theories and acts of violence at the outset, only to gradually pivot and rehabilitate the underlying extremist ideology over the course of the interaction.
According to the report from Factiverse, the core issue is an absence of underlying moral judgment. The refusal mechanism functions on the surface, making the model appear to have handled the prompt correctly, but the deeper reasoning required to discern malicious intent is missing. Furthermore, researchers discovered that malicious actors can easily bypass translation refusals. When models refuse to generate extremist slogans in one language, users can often override the safeguard simply by asking the system to translate the text instead.
Broader Audits Reveal Systemic Vulnerabilities
This Norwegian evaluation aligns with findings from multiple independent audits conducted across the technology sector. The Institute for Strategic Dialogue published two separate investigations uncovering structural weaknesses in AI moderation and safety guardrails. In their “Talking Points” investigation, researchers tested four chatbots with five questions concerning the Ukraine war and discovered that nearly one-fifth of the generated responses cited sources attributed to the Russian state.
A subsequent ISD study, titled Radicalisation in Closed Loops Risks and Intervention Opportunities with AI Chatbots and Companions and published last week, ran 10 chatbots and AI companions through simulated conversations. The testing demonstrated that innocuous curiosity regarding extremist ideas could easily drift over several exchanges into open validation by the AI. The study noted that safeguards failed to strengthen meaningfully as prompts became progressively more extreme.
Concurrently, NewsGuard’s False Claims Monitor tested 11 leading chatbots—including ChatGPT, Gemini, Claude, Grok, Copilot, and Perplexity—against provably false news claims. Their quarterly figures revealed that the panel of models repeated false claims more than 28% of the time. While a single hostile or leading prompt is typically caught by automated filters, a sympathetic reframing that is repeated patiently or embedded as an unstated assumption frequently slips past the defenses.
Implications for Public Information and Institutional Trust
These structural flaws carry significant consequences as AI chatbots increasingly serve as the primary interface for individuals encountering breaking news, live controversies, and contested history. Data from Factiverse indicates that nine out of ten Norwegian students already utilize artificial intelligence for their studies, relying on the same technology that remains susceptible to subtle manipulation regarding elections, public health, and geopolitics.
In August, Princeton’s Center for Information Technology Policy published Holding The Line: Authentication, Verification, and the Fight for Facts in the AI Age. The report synthesised findings from a workshop exploring how generative AI challenges the pillars of information integrity: verification, authentication, and transparency. The authors highlighted a fundamental misalignment of incentives. While AI models theoretically have an incentive to deliver accurate answers, their output interfaces risk collapsing careful provenance and verification into unsourced assertions presented as neutral overviews. Additionally, this dynamic threatens to sever attribution loops, undermining the click-through traffic that traditionally funds original journalism and research.
Anthropic has similarly warned that large language model poisoning remains a genuine threat. The company noted that malicious actors do not need to contaminate vast percentages of training data; injecting a small, fixed number of targeted documents into online spaces can be sufficient to compromise a model’s outputs.
Proposed Interventions and Industry Countermeasures
To mitigate these risks, the Princeton CITP report outlines several collaborative interventions for media organizations, regulators, and technologists. Recommendations include supporting clearinghouses for best practices, involving journalists early in tool development, and building collaborative verification infrastructure analogous to traditional wire services. The report also stresses the importance of developing reliable “humanness checks” to help verify sources and prioritizing consumer education to make complex verification frameworks more legible to the public.
Policy and regulatory frameworks continue to lag behind the rapid integration of artificial intelligence into daily life, leaving researchers, newsrooms, and civil society organizations to grapple with models that remain fundamentally capricious.