Authors
Mohan Tang and Sidi Lu
Context
This project was led by my student Mohan Tang, whom I mentored during my time at Tencent AI Lab Seattle.
Overview
Turbo Connection, or TurboConn, modifies a Transformer’s residual pathways so that higher-layer hidden states at token t can flow into lower layers at token t + 1. This creates a deeper effective computation path across generated tokens without requiring the full model to be trained again from scratch.
The work also studies connection density: dense routing between layers performs better than sparse alternatives that pass only a single hidden state or vector.
Results
Fine-tuning pretrained language models with TurboConn improves accuracy by 0.9 to more than 10 percentage points on GSM8K, Parity, and multi-step arithmetic. On the Parity task, it raises a fine-tuned Qwen3-1.7B model from 53.78% to 100% accuracy, with negligible impact on generation latency.
Paper
Status
Preprint, 2026.
Citation
Mohan Tang and Sidi Lu. “Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers.” arXiv:2602.17993, 2026.