• Login
  • Register

Work for a Member organization and need a Member Portal account? Register here with your official email address.

Project

Simulating "Neural Integration" of Human and AI through Mechanistic Interpretability as Design Provocation

Cyborg Psychology

Groups

As AI systems evolve from external tools to wearable interfaces and prospective neural implants, the proximity between human cognition and AI computation continues to diminish. This trajectory toward neural integration, where AI systems may directly interface with the human brain, presents unprecedented opportunities for cognitive augmentation alongside profound risks to human agency. However, investigating these dynamics empirically remains infeasible due to the nascent state of neural interface technology and significant ethical constraints.

We present a novel simulation that leverages mechanistic interpretability of large language models (LLMs) to model the functional dynamics of human-AI neural integration. Our approach is grounded in a key structural parallel: both biological brains and neural networks produce observable behaviors from internal distributed representations that can be decoded and systematically manipulated along interpretable dimensions. We introduce the concept of an AI-Symbiont, a hypothetical system capable of decoding and stimulating human neural activations, and simulate its interaction with a human cognitive proxy.

We employ two computational systems: a large language model serving as a proxy for a human cognitive system with interpretable internal activation states, and an AI-Symbiont module that reads and modulates these states through contrastive activation analysis and targeted activation steering. We demonstrate how integration might augment capabilities by applying steering vectors aligned with task demands, while revealing concerning amputation pathways where misaligned stimulation degrades performance.

Our results show that amputation produces asymmetrically larger degradation than augmentation produces improvement, suggesting the risk surface of misaligned neural stimulation may substantially exceed the benefit surface of aligned intervention. We further identify risk scenarios and provide HCI researchers with a computationally grounded testbed for anticipating design challenges and ethical implications of neural AI integration before such technologies are deployed at scale.