As AI systems evolve from external tools to wearable interfaces and prospective neural implants, the proximity between human cognition and AI computation continues to diminish. This trajectory toward neural integration, where AI systems may directly interface with the human brain, presents unprecedented opportunities for cognitive augmentation alongside profound risks to human agency. However, investigating these dynamics empirically remains infeasible due to the nascent state of neural interface technology and significant ethical constraints.
We present a novel simulation that leverages mechanistic interpretability of large language models (LLMs) to model the functional dynamics of human-AI neural integration. Our approach is grounded in a key structural parallel: both biological brains and neural networks produce observable behaviors from internal distributed representations that can be decoded and systematically manipulated along interpretable dimensions. We introduce the concept of an AI-Symbiont, a hypothetical system capable of decoding and stimulating human neural activations, and simulate its interaction with a human cognitive proxy.