Researchers have introduced Telescopic Language Models (TLM), a new architecture that allows a single Transformer model to adapt to various computational budgets. This approach trains a nested-capacity model using stochastic prefix supervision, enabling efficient serving across different performance needs.
A paper published on arXiv details the development of Telescopic Language Models (TLM), a novel approach to Transformer architecture designed to serve multiple computational budgets with a single model. Traditional methods require separate training or compression runs for each desired performance level. TLMs, however, are trained as a nested-capacity Transformer using stochastic prefix supervision. This method ensures that at every training step, a randomly truncated prefix of the model's capacity is trained against the full next-token target, alongside a full-capacity pass. The result is a single trained artifact that functions as a valid language model at every depth or capacity level. This training process involves only two forward-backward passes per step without requiring architectural changes or additional components at inference time.
The TLM approach contrasts with fixed-exit suites like Matryoshka Language Model Suites (MLMS), which occupy specific points in a design space. The researchers note that supervising only a few fixed exits in MLMS leaves the nested model performing at chance levels in other capacities. In contrast, a single TLM run serves as a valid language model across all its layer prefixes, demonstrating improved perplexity and performance on perplexity-sensitive downstream tasks. On a proxy suite with 200 million tokens, a single TLM run reduced the area under the quality-budget curve by 43-44% compared to fixed-exit suites, while matching their full-capacity performance at approximately 12% lower GPU cost per run. The density of prefix sampling during training is a tunable parameter. Concentrating this density on specific depths can recover fixed-exit quality at those points, effectively making operating points a training-time choice rather than an architectural one. This suggests that the training objective, not just the nested structure itself, is key to a model's elasticity.
The research paper, authored by Zhilin Guo and colleagues, highlights that the TLM training objective is what imparts elasticity to the model, rather than the nested architecture alone. This method allows for a continuum of model capacities from a single training run, offering a more efficient way to deploy language models that need to cater to diverse computational constraints. The findings suggest that TLMs can provide a more adaptable and cost-effective solution for serving language models across a range of performance requirements.
