Skip to content

Add vision support to LlamaLanguageModel via mtmd - #213

Open
james-333i wants to merge 2 commits into
huggingface:mainfrom
james-333i:feat/llama-vision
Open

Add vision support to LlamaLanguageModel via mtmd#213
james-333i wants to merge 2 commits into
huggingface:mainfrom
james-333i:feat/llama-vision

Conversation

@james-333i

Copy link
Copy Markdown
Contributor

Image segments threw unsupportedFeature because the backend had no multimodal path, even though the prebuilt llama.cpp binaries ship the mtmd library. This accepts an mmprojPath at initialization, replaces image segments with the mtmd media marker during prompt formatting, and evaluates text and image chunks through mtmd_helper_eval_chunks before sampling continues. Both respond and streaming support images, and models without a projector keep rejecting image input. Stacks on #195; the squash-merge will sort out the shared commit.

streamResponse consumed an inner AsyncThrowingStream whose builder ran
the entire generation loop synchronously on the consuming task, so
every snapshot buffered and arrived in one burst after generation
finished.

Yield snapshots directly from the generation loop on the streaming
task, and check for task cancellation between tokens so an abandoned
stream stops decoding promptly.
Image segments threw unsupportedFeature because the backend had no
multimodal path, even though the prebuilt llama.cpp binaries ship the
mtmd library and its helpers.

Accept an mmprojPath at initialization and load the projector next to
the model. When a projector is present, prompt formatting replaces
each image segment with the mtmd media marker and collects payloads in
order, then generation tokenizes the marker-annotated prompt with
mtmd_tokenize and evaluates text and image chunks through
mtmd_helper_eval_chunks before sampling continues from the resulting
position. Both respond and streaming support images, and models
without a projector keep rejecting image input.

Adds live tests generating from an embedded test image through both
paths.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant