Original sourceBlockchain.News
Summary
Anthropic released the GRAM (Gradient Routing Assisted Module) method, which controls AI dual-use knowledge without extensive retraining by adding dedicated neuron modules within Transformer layers. During training, only relevant modules are updated; afterward, specific capabilities (such as cybers…
Key points
- AI safety researchers and model deployers can more precisely control high-risk model knowledge, reducing misuse risk while maintaining excellent performance on everyday tasks.
- GRAM offers a pluggable safety mechanism that allows finer-grained model capability control, influencing risk management strategies for both open-source and commercial AI models.
- When deploying large language models, developers can selectively disable potentially harmful modules, lowering compliance and ethical risks for safer AI applications.
- Unlike traditional refusal responses or classifier-based filtering, GRAM manages knowledge directly at the model architecture level, akin to designing on-off switches for software features.
Editorial note
This page is Code & Chain's editorial summary of public sources. It may be prepared with AI assistance and published through an automated workflow. Refer to the original sources; this content is not investment, legal, or tax advice.