Files
2025-08-22 17:04:24 +02:00

1.3 KiB

🔠 Tiktoken Tokenizer Info

This node takes text as input, and returns a bunch of data from the tiktoken tokenizer.

It returns the following values:

  • token_count: Total number of tokens
  • character_count: Total number of characters
  • word_count: Total number of words
  • split_string: Tokenized list of strings
  • split_string_list: Tokenized list of strings (output as list)
  • split_token_ids: List of token IDs
  • split_token_ids_list: List of token IDs (output as list)
  • text_hash: Text hash
  • special_tokens_used: Special tokens used
  • special_tokens_used_list: Special tokens used (output as list)
  • token_chunk_by_size: Returns the input text, split into different strings in a list by the token_chunk_size value.
  • token_chunk_by_size_to_word: Same as above but respects "words" by stripping backwards to the nearest space and splitting the chunk there.
  • token_chunk_by_size_to_section: Same as above, but strips backwards to the nearest newline, period or comma.

Tokenization & Word Count example

image


Chunking Example

image