(github.com)

279 points matthewolfe | 1 comments | 30 Jun 25 12:33 UTC | HN request time: 0.001s | source

TokenDagger is a drop-in replacement for OpenAI’s Tiktoken (the tokenizer behind Llama 3, Mistral, GPT-3.*, etc.). It’s written in C++ 17 with thin Python bindings, keeps the exact same BPE vocab/special-token rules, and focuses on raw speed.

I’m teaching myself LLM internals by re-implementing the stack from first principles. Profiling TikToken’s Python/Rust implementation showed a lot of time was spent doing regex matching. Most of my perf gains come from a) using a faster jit-compiled regex engine; and b) simplifying the algorithm to forego regex matching special tokens at all.

Benchmarking code is included. Notable results show: - 4x faster code sample tokenization on a single thread. - 2-3x higher throughput when tested on a 1GB natural language text file.

Show context

npalli ◴[30 Jun 25 13:13 UTC] No.44422888[source]▶

>>44422480 (OP) #

Kudos, I think (in the short term at least) there is a large amount of perf. optimization to be found by coding parts of the whole AI/ML infrastructure in C++ like this one, not as a rewrite (god no!) but drop in and fix key bottlenecks. Anytime I see someone (seems Chinese engineers are good at this) put something out in C++, good chance some solid engineering tradeoffs have been made and dramatic improvement will be seen.

replies(4): >>44424382 #>>44424572 #>>44424990 #>>44427963 #

notatallshaw ◴[30 Jun 25 21:15 UTC] No.44427963[source]▶

>>44422888 #

It looks like TikToken is written in Rust (https://github.com/openai/tiktoken/tree/main/src), are the gains here actually from porting to C++?

replies(1): >>44430216 #

1. fhub ◴[01 Jul 25 03:11 UTC] No.44430216[source]▶

>>44427963 #

From the post

Profiling TikToken’s Python/Rust implementation showed a lot of time was spent doing regex matching. Most of my perf gains come from a) using a faster jit-compiled regex engine; and b) simplifying the algorithm to forego regex matching special tokens at all.

↑

Show HN: TokenDagger – A tokenizer faster than OpenAI's Tiktoken