1. MiniMax-H3 by MiniMaxAI
image-text-to-video · diffusers · Likes: 3574 · Downloads: 59368 · View on Hugging Face →
Tags: diffusers, safetensors, text-to-video, image-to-video, image-text-to-video, video-to-video, text-to-audio-video
What it is. MiniMax-H3 is a multimodal model in the diffusers library whose primary pipeline is image-text-to-video, with tags also covering synchronized audio-video generation from text, image, video, audio, and reference inputs.
Why it's trending. It is trending this week with a trending score of 1346, backed by strong community engagement of 3,574 likes and over 59,000 downloads.
Who should use it. Developers and researchers building multimodal generation pipelines who need to produce synchronized audio-video outputs from combinations of text, images, video, audio, or reference inputs.
2. Muse-Glimmer-30B by meta-models
image-text-to-text · transformers · Likes: 1097 · Downloads: 0 · View on Hugging Face →
Tags: transformers, safetensors, muse_glimmer, image-text-to-text, conversational, arxiv:2504.13181, arxiv:2602.06036
What it is. Muse-Glimmer-30B is a 30-billion parameter multimodal model from meta-models designed for image-text-to-text generation and conversational tasks.
Why it's trending. It is trending this week due to strong community interest, evidenced by over 1,000 likes and a high trending score, paired with its open Apache 2.0 license and linked associated research papers for the image-text-to-text task.
Who should use it. Developers and researchers building multimodal conversational applications that process combined image and text inputs are the target audience for this model.
3. DeepSeek-V4-Flash-0731 by deepseek-ai
text-generation · transformers · Likes: 3150 · Downloads: 1048685 · View on Hugging Face →
Tags: transformers, safetensors, deepseek_v4, text-generation, conversational, arxiv:2606.19348, license:mit
What it is. DeepSeek-V4-Flash-0731 is a text-generation model in the conversational modality, built on the transformers library and offered in fp8 and 8-bit quantized formats.
Why it's trending. It has surpassed one million downloads and earned 3,150 likes with a trending score of 668, indicating strong and rapid community adoption this week.
Who should use it. Developers and teams seeking a deployable, quantized text-generation model for conversational applications—especially on Azure endpoints—should consider this model.
4. MiniMax-H3-Turbo-Lora by larryvrh
text-to-video · Likes: 651 · Downloads: 0 · View on Hugging Face →
Tags: text-to-video, text-to-audio, audio-video, lora, minimax-h3, comfyui, base_model:Comfy-Org/MiniMax-H3
What it is. It is a LoRA adapter for the MiniMax-H3 model that performs text-to-video generation with associated audio-video capabilities via ComfyUI.
Why it's trending. It is notable this week for drawing strong community engagement, with 651 likes and a high trending score of 612 despite zero recorded downloads.
Who should use it. It is intended for creators and developers working in ComfyUI who want to use a MiniMax-H3-based LoRA for text-to-video and text-to-audio workflows.
5. MiniMax-H3 by Comfy-Org
Model · diffusion-single-file · Likes: 1212 · Downloads: 6798796 · View on Hugging Face →
Tags: diffusion-single-file, comfyui, base_model:MiniMaxAI/MiniMax-H3, base_model:finetune:MiniMaxAI/MiniMax-H3
What it is. MiniMax-H3 is a single-file diffusion model developed by Comfy-Org for use with the ComfyUI interface to perform image generation tasks.
Why it's trending. It currently holds a Hugging Face trending score of 535, alongside over 6.7 million downloads and more than 1,200 community likes, reflecting strong recent user adoption.
Who should use it. It is suited for ComfyUI users seeking a portable, low-friction diffusion model to run image generation workflows without extensive setup.
6. Kimi-K3 by moonshotai
image-text-to-text · transformers · Likes: 10526 · Downloads: 1565484 · View on Hugging Face →
Tags: transformers, safetensors, kimi_k3, feature-extraction, compressed-tensors, conversational, image-text-to-text
What it is. Kimi-K3 is a multimodal image-text-to-text model developed by moonshotai that accepts both image and text inputs to generate text outputs.
Why it's trending. It is trending this week due to its high Hugging Face user engagement, with over 1.56 million downloads, more than 10,500 likes, and a trending score of 385.
Who should use it. Developers building multimodal text generation applications that combine image and text inputs, including conversational and visual content analysis use cases, should use this model.
7. LFM2.5-2.6B by LiquidAI
text-generation · transformers · Likes: 550 · Downloads: 93668 · View on Hugging Face →
Tags: transformers, safetensors, lfm2, text-generation, liquid, lfm2.5, edge, conversational, ar, zh, en, fr, de, hi, id
What it is. LFM2.5-2.6B is a 2.6B-parameter text-generation model from LiquidAI, fine-tuned from LFM2.5-2.6B-Base for conversational use across many languages.
Why it's trending. It is gaining attention this week with a high trending score of 373, 550 likes, and over 93,000 downloads, reflecting strong interest in its edge-focused, multilingual design.
Who should use it. Developers and researchers building multilingual conversational AI applications, especially those targeting edge deployment scenarios.
8. Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF by DavidAU
image-text-to-text · Likes: 1898 · Downloads: 2521093 · View on Hugging Face →
Tags: gguf, unsloth, fine tune, heretic, uncensored, abliterated, ara, MTP GGUF Quants, Regular GGUF Quants, qwen3.6
What it is. A GGUF-quantized 27B-parameter fine-tuned language model based on Qwen3.6, configured for image-text-to-text conversation, reasoning, creative writing, and coding tasks.
Why it's trending. It has accumulated over 2.5 million downloads and nearly 1,900 likes alongside a high trending score, reflecting strong community interest in its multi-stage uncensored fine-tune of the Qwen3.6 base.
Who should use it. Creative writers, roleplayers, and developers looking for an unrestricted, multilingual conversational model with reasoning and image-text understanding.
9. Minimax-h3-Turbo by lightx2v
image-to-video · diffusers · Likes: 341 · Downloads: 20376 · View on Hugging Face →
Tags: diffusers, t2v, i2v, r2v, image-to-video, en, zh, base_model:MiniMaxAI/MiniMax-H3
What it is. Minimax-h3-Turbo is a Diffusers-based model tagged for image-to-video generation (with additional text-to-video and reference-to-video tags) that is a finetune of the MiniMax-H3 base model.
Why it's trending. It is gaining traction this week with a trending score of 327, supported by 341 likes and more than 20,000 downloads on the Hugging Face Hub.
Who should use it. Developers and researchers using the Diffusers library who need an image-conditioned video generation model derived from the MiniMax-H3 base.
10. Muse-Glimmer-30B-GGUF by unsloth
image-text-to-text · transformers · Likes: 305 · Downloads: 0 · View on Hugging Face →
Tags: transformers, gguf, unsloth, meta, image-text-to-text, arxiv:2504.13181, arxiv:2602.06036
What it is. A GGUF-quantized 30B vision-language model that performs image-and-text to text generation, derived from Meta's Muse-Glimmer-30B base model and published by Unsloth under the Apache-2.0 license.
Why it's trending. It is climbing the Hugging Face trending charts this week with a high trending score of 301 and 305 likes, driven by community interest in Unsloth's GGUF build of a 30B image-text-to-text model.
Who should use it. Developers and hobbyists who want to run a quantized, locally-deployable multimodal model that accepts images plus text and produces conversational text responses.