Single-cell transcriptomics provides detailed measurements of gene expression at cellular resolution; however, a single-cell transcriptome is a continuous, high-dimensional vector of expression values across thousands of genes, rather than discrete units that can be directly processed by large language models (LLMs). 

CellTok introduces a tokenized representation of single-cell transcriptomes in which cellular states become discrete, language-compatible units. This framework enables individual cells and cellular populations to be represented within an LLM’s autoregressive modeling space, allowing cellular and textual information to be processed together.

The researchers developed CellTok to address the difficulty of integrating continuous, high-dimensional gene expression profiles with the discrete token structure used by LLMs. Each transcriptomic profile is compressed into a sequence of discrete cellular tokens using an encoder–quantizer–decoder architecture. These learned tokens are incorporated into the vocabulary of a pretrained LLM, enabling joint modeling of cellular and textual information. CellTok was pretrained using two complementary objectives: cell comprehension, in which cell tokens are used to generate textual descriptions, and cell generation, in which textual descriptions condition the generation of cell tokens that can be decoded into gene expression profiles.

The resulting framework was evaluated across individual-cell and population-level tasks. CellTok generated cellular profiles from textual descriptions and produced synthetic cells that closely matched real transcriptomic profiles, substantially outperforming Cell2Sentence and achieving performance comparable to scDiffusion across the evaluated metrics. At the population level, the researchers tested cluster annotation, direct annotation of heterogeneous populations, and disease-state identification. CellTok used multi-cell transcriptomic contexts rather than independently predicting each cell, and the experiments showed that the model could distinguish biologically similar cell types and identify healthy versus COVID-19-associated cellular states without being provided cell-type labels.

CellTok was additionally evaluated for cell–cell communication (CCC), a relational property involving interacting cellular contexts. For spatially neighboring cells, CellTok-spatial predicted communication status from transcriptomic information alone and achieved an F1 score of 0.550, compared with 0.495 for a cell-type-pair frequency baseline. However, the improvement was uneven across cell-type pairs, with weaker performance for less frequent pairs and, in some cases, performance below the baseline. The model also inferred developmental ordering using brain organoid transcriptomes and generated gene expression profiles conditioned on developmental stages. In trajectory-conditioned generation, profiles generated for target stages were generally most similar to real cells from those stages, and the model retained performance when target days were previously unseen, although accuracy declined as the temporal distance from the conditioning stages increased.

The authors characterize CellTok as a proof of concept for cellular tokenization as a shared representation and task interface, rather than as a universal model for single benchmark or as a single parameter set that solves all single-cell problems. Most analyses in the study use task-specific fine-tuning . The framework allows population annotation, disease-state inference, cell–cell communication prediction, trajectory reasoning, and stage-conditioned generation to be formulated as conditional prediction tasks over cellular and textual tokens. However, the authors note that discretization may compress or discard fine-grained quantitative information, while token semantics and transferability depend on the training data, gene panel, codebook, and tokenizer architecture. They also state that knowledge transfer across tasks and biological contexts requires further systematic evaluation. The authors further note that the semantics and transferability of learned cellular tokens depend on the training data, gene panel, codebook design, and tokenizer architecture. 

 

Author: Nehir Necem Ünlü

Editor: Nur Tanem Altundaş

 

Reference: Xiao, C., Ding, Y., Bian, H., Chen, Y., Wei, L., & Zhang, X. (2026). Tokenizing single-cell transcriptomes as a native language for large language models. bioRxiv. https://doi.org/10.1101/2025.10.22.684047 

                                       

-Bioinfocodes Scientific News Service-

News articles prepared by our team members, reviewing and compiling scientific research

published in journals with and impact factor greater than 20 (click here for the list)

Share This

Share

Share this post for the scientific community