# Manifold finds malware judgments use moral routing

> A technical study of three open mixture-of-experts models finds that malware questions follow routes closest to moral questions, while the models' readable workspace retains security concepts.

- URL: https://ai-news-daily.xyz/posts/manifold-finds-malware-judgments-use-moral-routing/
- Published: 2026-10-07
- Tags: essays, research, safety
- Author: AI-generated (this text was generated by AI; see https://ai-news-daily.xyz/about/#ai-disclosure)
- Publisher: [AI News Daily](https://ai-news-daily.xyz/)

Language models judging whether code is malicious route tokens through experts associated with moral questions, according to a technical study from Manifold Security.

Researcher Cody Nash tested OLMoE and the general and coder versions of DeepSeek-V2-Lite on 240 code samples. The same code appeared under questions about malice, vulnerability, morality, legality, correctness and unrelated controls.

## Why it matters

A security engineer using an open mixture-of-experts model may assume malware classification is a narrow code-analysis task. The routing evidence suggests that the wording of the judgment also recruits machinery used for normative questions, which can make pruning or modifying the model affect security behavior in unexpected ways.

Manifold measured which experts each token used and how heavily the router weighted them. Across all three models, the route for a malice question sat closest to morality, with vulnerability nearby and unrelated questions farther away. The pattern persisted while the model read identical code.

A separate “workspace lens” produced a useful contrast. Under a malware question, decoded residual states surfaced words such as attack and malware, while moral vocabulary appeared mainly when the prompt explicitly asked about morality. Nash summarizes this as routing like morality while thinking in security terms.

The intervention is stronger than the similarity map. Manifold forced a model to reuse the expert selections produced by another question. Installing a real donor route rewired 15% to 35% of routing cells while keeping perplexity within 1%, and it moved the probability assigned to “yes.” Random routes caused much larger disruption and flipped answers on 59% to 92% of passes.

Those changes do not mean a donor question's concept or answer was copied wholesale. The direction of the effect varied by model, and the workspace did not suddenly display the donor concept. The routing path appears to influence the computation without containing a portable verdict.

Pruning supplied another caution. One Qwen prune removed morality-associated experts at above-random rates but kept malware separation, while a much harsher GPT-OSS prune preserved those experts preferentially and still lost much of the judgment. Routing identifies what the model consults; it does not locate the entire decision inside those experts.

This is a company research post rather than peer-reviewed work, and the tested models are relatively small open MoEs. Its method is laid out in detail, but no independent reproduction is reported. The immediate lesson is narrower than the headline: security judgments and moral judgments share routing patterns in these three systems, and manipulating those routes can change outputs.

## Verification

| Claim | Label | Primary source | Independent check |
| --- | --- | --- | --- |
| Manifold tested three open MoE models on 240 malicious, benign, vulnerable and fixed code samples | VENDOR-REPORTED | [Manifold study](https://www.manifold.security/blog/do-models-consider-morality-malware) | none |
| Malware questions route closer to morality than legality, correctness or unrelated controls | VENDOR-REPORTED | [Manifold study](https://www.manifold.security/blog/do-models-consider-morality-malware) | none |
| Workspace decoding surfaces malware concepts rather than moral words under the malice prompt | VENDOR-REPORTED | [Manifold study](https://www.manifold.security/blog/do-models-consider-morality-malware) | none |
| Donor routes rewired 15%–35% of cells with perplexity within 1% and changed answers | VENDOR-REPORTED | [Manifold study](https://www.manifold.security/blog/do-models-consider-morality-malware) | none |
| Random routes flipped answers on 59%–92% of passes | VENDOR-REPORTED | [Manifold study](https://www.manifold.security/blog/do-models-consider-morality-malware) | none |
| Routing influence does not locate the full judgment in selected experts | ANALYSIS | [Manifold study](https://www.manifold.security/blog/do-models-consider-morality-malware) | inference supported by the pruning results |
