Add SmolLM2 support (1.7B, 360M, 135M) - #175
Open
stikves wants to merge 5 commits into
Open
Conversation
Add HuggingFace SmolLM2 model family for macOS and iOS export: - 1.7B: INT4 quantized with FP16 embedding (tied weights) - 360M/135M: FP16 (small enough uncompressed) SmolLM2 uses the Qwen2 architecture under a different model_type. Adds _model_type_override field to ModelPreset to route export without registering a global type mapping. Custom quantization config keeps embedding/lm_head at FP16 because SmolLM2 ties these weights — INT4 on the shared tensor degrades long-form generation quality (+26% PPL vs +15% with FP16 embedding).
stikves
requested review from
Lewis300,
carinapeng,
kevchengcodes,
pkmandke and
tjia1818
August 15, 2026 16:48
transformers 5.x requires rope_parameters to be a dict (not None) when initializing Qwen2RotaryEmbedding. Pass rope_parameters and rope_theta directly in the config constructor instead of setting rope_scaling=None post-init.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add HuggingFace SmolLM2 model family for macOS and iOS export:
SmolLM2 uses the Qwen2 architecture under a different model_type. Adds _model_type_override field to ModelPreset to route export without registering a global type mapping.
Custom quantization config keeps embedding/lm_head at FP16 because SmolLM2 ties these weights — INT4 on the shared tensor degrades long-form generation quality (+26% PPL vs +15% with FP16 embedding).