A research paper published in Nature on October 7 describes a method for converting existing language models to process text as bytes. The method, called byteification, changes the model’s input and output pathway. It does not simply remove every internal text unit. Instead, it groups byte sequences into variable-length patches before the main transformer processes them.
200 Ways To Make Money With AI
Turns out, AI is good for more than writing emails and generating LinkedIn posts. This guide is packed with 200+ actionable ways to build income with AI.
Inside you'll discover:
200 curated AI income ideas for beginners and pros alike
Easy-to-start opportunities you can launch this week
Real world applications for today’s top AI tools
AI powered business models built for today’s economy
Creative ways to turn trends into revenue
AI is changing how people work and how people earn, far beyond the simple “write me an email” prompt. Discover the possibilities and download the free guide to start cashing in today.
Most common language models use subword vocabularies. A tokenizer breaks text into pieces that may be whole words, word fragments, or punctuation. This approach is efficient for ordinary writing, but can make exact character-level work harder. Counting letters, handling unusual spellings, and reading uncommon strings depend on details that a subword representation may not expose cleanly.
The paper’s architecture tries to preserve a capable model’s learned knowledge while changing how text reaches it. An encoder first represents individual bytes. A boundary predictor groups them into patches, which pass through the model’s larger transformer. A decoder then returns the output to a byte sequence. The arrangement retains an internal compression stage, rather than requiring the main transformer to handle every byte as a separate step.
The researchers describe two training stages. First, new local components learn to reproduce the source model’s behavior while its central transformer stays frozen. In the second stage, the full system is trained to use byte-level information. The paper applies the approach to OLMo, Qwen3, and Llama models at one to eight billion parameters. This requires additional training; byteification is not a no-data or no-computation change.
Learn AI in 5 minutes a day
You don't have to scroll every AI thread, track every new tool, or watch every demo.
The Rundown AI breaks it all down for you — the latest AI news, tools, and tutorials in one free 5-minute email every morning.
Trusted by 2M+ professionals at Apple, Google, and NASA.
The paper reports byteified versions of Olmo, Qwen, and Llama models at roughly one to eight billion parameters. The authors compare them with their original subword models and with earlier byte-level systems. They report that the converted models approach their source models across several standard evaluations, while gaining on selected tests that examine character-level understanding. The results concern those models and test sets, not every language model or task.
This distinction matters for practical use. Exact spellings and rare strings can contain details that subword units do not expose cleanly. A byte-level pathway may give a model more direct access to those details. Yet performance depends on patch formation, model size, training data, and the task. It does not follow that every byte-based model will be faster, more accurate, or easier to deploy.
The technical trade-off is not simply “characters instead of tokens.” Byteification still learns patches inside the model. A large transformer then handles global context. The authors’ goal is to replace fixed external subword boundaries with boundaries the model can adapt. That may improve flexibility while retaining efficient processing. It also means the method’s success depends on both local byte components and the pretrained model they surround.
The publication is a research result, not evidence of broad deployment. The work tests a limited set of model families and benchmark tasks. Its comparisons cannot establish equal gains across scripts, languages, or live applications. The method also requires substantial training and model engineering, even when that cost is smaller than training a comparable system from scratch.
The useful change is methodological. Researchers can now test whether a model’s text representation, rather than only its scale, limits some exact-string tasks. Follow-up comparisons will need to report language coverage, compute, throughput, and failures alongside benchmark scores. For this week, byteification is a concrete research direction, with measurable benefits in selected tests and important questions still open.
Your SOC 2, handled end to end with AI agents.

Enterprise buyers won’t put your product near their customer data without a SOC 2 report. Twenty-one state privacy laws now hold them accountable for the vendors they share data with, so their diligence lands on you.
Sprinto gets you audit ready in 14 days, across three working sessions. AI agents connect your stack, collect the evidence auditors ask for, and close the gaps as they appear. You approve, they execute.
Sprinto also answers the security questionnaires your buyers send, using the same evidence base, so security reviews stop holding up deals.
No compliance hire. No consultant. Your auditor signs off.




