1. MiniMax-H3 by MiniMaxAI

image-text-to-video · diffusers · Likes: 2891 · Downloads: 18112 · View on Hugging Face →

Tags: diffusers, safetensors, text-to-video, image-to-video, image-text-to-video, video-to-video, text-to-audio-video

What it is. MiniMax-H3 is a multimodal image-text-to-video generation model that produces synchronized audio-video content from text, image, video, or audio inputs.

Why it's trending. It has garnered 2,891 likes and 18,112 downloads with a high trending score of 2,723, reflecting strong community interest in its broad multimodal audio-video generation capabilities.

Who should use it. Developers and creators building pipelines for text-to-video, image-to-video, or synchronized audio-video generation using the Diffusers library should consider this model.

2. DeepSeek-V4-Flash-0731 by deepseek-ai

text-generation · transformers · Likes: 2721 · Downloads: 702709 · View on Hugging Face →

Tags: transformers, safetensors, deepseek_v4, text-generation, conversational, arxiv:2606.19348, license:mit

What it is. DeepSeek-V4-Flash-0731 is a transformer-based text-generation model from DeepSeek tagged for conversational language tasks and available under an MIT license.

Why it's trending. It is trending on the Hugging Face Hub this week with a high trending score, strong engagement signals of roughly 702,709 downloads and 2,721 likes, and recognition as a recent DeepSeek V4 Flash release supporting 8-bit and fp8 quantization with Azure deployment.

Who should use it. Developers and researchers building conversational AI applications or text-generation pipelines who want an MIT-licensed, quantization-friendly model deployable on cloud endpoints.

3. MiniMax-H3 by Comfy-Org

Model · diffusion-single-file · Likes: 907 · Downloads: 3139920 · View on Hugging Face →

Tags: diffusion-single-file, comfyui, base_model:MiniMaxAI/MiniMax-H3, base_model:finetune:MiniMaxAI/MiniMax-H3

What it is. A single-file diffusion checkpoint packaged by Comfy-Org for ComfyUI, based on the MiniMaxAI MiniMax-H3 base model for image generation workflows.

Why it's trending. It shows notable weekly traction this period, reflected in a high trending score alongside millions of downloads and several hundred likes, signaling strong community uptake of the ComfyUI-ready MiniMax-H3 build.

Who should use it. ComfyUI users and visual creators who want a plug-and-play diffusion checkpoint derived from the MiniMax-H3 base model for their image generation pipelines.

4. Kimi-K3 by moonshotai

image-text-to-text · transformers · Likes: 10257 · Downloads: 1308186 · View on Hugging Face →

Tags: transformers, safetensors, kimi_k3, feature-extraction, compressed-tensors, conversational, image-text-to-text

What it is. Kimi-K3 is a multimodal image-text-to-text model from moonshotai designed to process and generate text from combined image and text inputs.

Why it's trending. It is trending this week, backed by a high trending score of 791, over 1.3 million downloads, and more than 10,000 likes that reflect strong community traction.

Who should use it. AI developers, researchers, and teams building vision-language applications that require processing and generating text from combined image and text inputs are the intended users of this model.

5. DeepSeek-V4-Flash-0731-GGUF by unsloth

Model · Likes: 577 · Downloads: 161253 · View on Hugging Face →

Tags: gguf, unsloth, deepseek_v4, deepseek, arxiv:2606.19348, base_model:deepseek-ai/DeepSeek-V4-Flash-0731

What it is. This is a quantized GGUF-format conversational large language model based on the DeepSeek-V4-Flash-0731 base model, released by Unsloth.

Why it's trending. It is trending this week due to strong user adoption, reflected in over 161,000 downloads, 577 likes, and a high trending score, as an endpoint-compatible, imatrix-optimized quantized variant of the DeepSeek-V4-Flash-0731 model.

Who should use it. It is intended for developers and AI enthusiasts looking to run or deploy a compact, endpoint-compatible version of DeepSeek-V4-Flash-0731 for local or cloud-based conversational text generation tasks.