AutoTokenizer tries to load the model config to detect tokenizer type,
which fails because transformers 5.x CONFIG_MAPPING doesn't know
florence2_language. Use BartTokenizerFast directly with tokenizer.json
fallback for models that use different tokenizer formats.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
When AutoTokenizer fails (e.g. Florence-2-Flux-Large), fall back to
loading tokenizer.json directly via the tokenizers library and wrap
it with PreTrainedTokenizerFast, extracting special tokens from
tokenizer_config.json manually.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Some Florence2 variants (e.g. Florence-2-Flux-Large) have added_tokens
in a format that fast tokenizers in transformers 5.x cannot parse.
Fall back to slow tokenizer on TypeError.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Some Florence2 variants (e.g. Florence-2-Flux-Large) use a RoBERTa
tokenizer. AutoTokenizer auto-detects the correct tokenizer class.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The florence2_models directory needs an __init__.py to be a proper
Python package, and imports from py/florence2_ultra.py need to use
relative imports (..florence2_models) since it's a sibling subpackage.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Uses local Florence2 modules with accelerate-based weight loading and
a custom Florence2Processor for transformers >= 5.0.0, while keeping
the existing loading path for older versions.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
In transformers 5.x, forced_bos_token_id was moved from PretrainedConfig
to GenerationConfig. Florence2's custom config accesses this attribute
during init, causing an AttributeError that prevents model loading and
leads to a NoneType crash when the processor is later called.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The old code used list.extend() on the outer (image-index) dimension,causing bboxes from different inputs to be treated as separate images rather than merged into the same image's bbox list.
transformers 5.x introduced two breaking changes that affect the
bundled GroundingDINO BertModelWarper:
1. `BertModel.get_head_mask()` was removed. Added a standalone
`_get_head_mask()` fallback with a `hasattr` check so both
transformers 4.x and 5.x work.
2. `BertModel.get_extended_attention_mask()` changed its 3rd
positional argument from `device` to `dtype`. The old call-site
passes a `torch.device`, which newer transformers interprets as
`dtype=torch.device` → TypeError. Replaced with a standalone
`_get_extended_attention_mask()` static method that also handles
PyTorch 2.9's prohibition of `1.0 - bool_tensor` by explicitly
casting to float32 first.
Fixes#140
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
fixes RuntimeError: Attempting to deserialize object on a CUDA device but torch.cuda.is_available() is False. If you are running on a CPU-only machine, please use torch.load with map_location=torch.device('cpu') to map your storages to the CPU.