source avatarMinty

Share

Could faster tokenization remove a major bottleneck in training large AI models? Gigatoken is an open-source tokenizer designed to make this process much faster by reading training data directly from files and distributing the work across multiple CPU cores. Before raw text can be used for training, it has to be converted into tokens that the model can process. This usually doesn’t matter for small datasets, but tokenizing larger datasets requires more time and CPU resources. Gigatoken processed roughly 2.7 billion GPT-2 tokens in about half a second on a high end server, which was nearly 1,000x faster than Hugging Face’s tokenizer in the project’s benchmark. For large training labs and data providers, that could mean fewer CPUs dedicated to preparing data and faster iteration whenever a dataset or tokenizer changes.

No.0 picture
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.