Skip to content
Pardeep Kaushik.

LLMs & Model Architecture

How AI Tokens Work: BPE Tokenization, Calculation and Practical Rules

A deep dive into Byte-Pair Encoding (BPE) and tokenization algorithms. Learn how tokenizers split text, handle whitespace, and impact LLM performance.

  • Tokenization
  • BPE
  • Machine Learning
  • LLM
  • Algorithms

The Mathematics of Text Ingestion

In our foundational guide on what are tokens in AI, we established that tokens are the fundamental building blocks of Large Language Models.

To software developers building AI-driven web apps and API integrations, a deeper question arises: How does a tokenizer actually decide where to split a word, and why does this matter for model accuracy and speed?

This guide explores the mechanics of Byte-Pair Encoding (BPE), tokenizer vocabulary design, and practical rules for optimizing token efficiency in software engineering.

The Byte-Pair Encoding (BPE) Algorithm Explained

Byte-Pair Encoding (BPE) was originally created in 1994 as a data compression algorithm. In modern AI, it was repurposed to build efficient vocabularies for language models.

How BPE Builds a Vocabulary:

  1. Base Characters: Start with all individual characters (letters, numbers, basic punctuation).
  2. Frequency Analysis: Scan a massive corpus of text and count which pairs of characters appear next to each other most frequently.
  3. Merge Step: Combine the most frequent pair into a new single token (for example, 't' + 'h' becomes 'th').
  4. Iterate: Repeat this merging process thousands of times until the target vocabulary size (typically 50,000 to 128,000 unique tokens) is reached.

As a result, frequent words like "the", "is", and "function" become single tokens, while rare words are split into known sub-word components.

Tokenizer Variations Across Frontier Models

Different AI providers train custom tokenizers with distinct vocabulary sizes:

Model FamilyTokenizer LibraryVocabulary SizeKey Optimization
OpenAI (GPT-4o)o200k_base (tiktoken)~200,000 tokensMassive multilingual & code efficiency
Anthropic (Claude)Custom Claude Tokenizer~100,000 tokensOptimized for English prose, XML, and code
Meta (Llama 3)Tiktoken-based BPE~128,000 tokensEnhanced multilingual and mathematical token balance

A larger vocabulary allows a model to compress more text into fewer tokens, directly reducing latency and cost per word.

Practical Example: Tokenizing Code with Tiktoken

In a Node.js or Python backend, calculating tokens before calling an AI model is straightforward:

# Python token counting example
import tiktoken

encoding = tiktoken.get_encoding("o200k_base")
text = "function calculateTotal(items) { return items.reduce((a, b) => a + b, 0); }"

tokens = encoding.encode(text)
print(f"Token count: {len(tokens)}")
# Outputs the exact integer token count for API budget planning

3 Practical Rules for Developers

  1. Minify JSON Payloads: When passing structured data to an LLM, strip unnecessary whitespace and indentation. Formatting a JSON payload with 4-space indentation can double its token consumption compared to a compact single-line JSON string.
  2. Standardize Variable Names: Use common naming conventions. Exotic camelCase concatenations (mySuperDuperCustomHelperFunction) can split into 6+ tokens, whereas standard names (calculateMetrics) consume fewer tokens.
  3. Use Prompt Caching: When utilizing models like Claude 3.7 via API, leverage Prompt Caching to avoid paying full price for repeatedly processed system instructions. Read more in our guide on Claude Code usage limits and pricing.

Need high-performance API architectures built for your web application? Explore our full-stack web development services or view our client reviews.

Frequently asked questions

What algorithm do modern LLMs use for tokenization?

Most modern Large Language Models (including GPT-4o, Claude, and Llama) use Byte-Pair Encoding (BPE) or variants like WordPiece and Unigram tokenization.

How does Byte-Pair Encoding (BPE) work?

BPE starts with individual characters and iteratively merges the most frequently occurring character pairs in a training dataset into single tokens until a predetermined vocabulary size (e.g., 100,000 tokens) is reached.

Why do numbers and math often confuse tokenizers?

Tokenizers often split multi-digit numbers unpredictably (e.g., '1024' might become '10' and '24' while '1025' is split as '1' and '025'), making basic arithmetic harder for language models unless trained on explicit digit-level tokens.

Can I inspect token counts before sending an API request?

Yes. Libraries like OpenAI's tiktoken or Anthropic's token counting endpoints allow developers to calculate exact token sizes locally before making expensive API calls.

About the author

Author

Pardeep Kaushik

Full Stack, WordPress & Shopify Developer

Pardeep Kaushik is a freelance Full Stack, WordPress and Shopify developer with 5+ years of experience building business websites, ecommerce stores and custom web applications. His work includes WordPress, WooCommerce, Elementor, Shopify, Liquid, React, Next.js, Node.js, AI integrations, APIs and production deployment.