This AI Model Has Native Text, Image, and Video Capabilities: Here’s What You Should Know

AI Video & Visuals


Overview

Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-GGUF is a locally runnable, text-to-text GGUF release with native text, image, and video capabilities. LuffyTheFox built it from the uncensored HauhauCS Qwen3.6-35B-A3B base, applied the Genesis tensor-repair process, and transferred about 2,000 blocks from two FFN expert tensors in a Hermes fine-tune to add Hermes-agent behavior. It has 35 billion total parameters, about 3 billion active per forward pass, a 262K-token native context window, and a hybrid MoE architecture combining Gated DeltaNet linear attention with full softmax attention.

The key practical point: this is a large GGUF model, not a lightweight 3B model; the active-parameter count does not describe the full storage or memory footprint. It runs in llama.cpp, LM Studio, koboldcpp, and other GGUF-compatible runtimes. The model card recommends NVFP4 or APEX quantization, and APEX Compact for 8 GB and 12 GB GPUs. The listed license is Apache-2.0, but check the terms of the upstream components before commercial deployment.

Best Use Cases

Local coding and precise instruction-following. The model card provides a dedicated thinking-mode sampling profile for coding and precise tasks: temperature 0.6, top_p 0.95, top_k 20, min_p 0.01, seed 42, with presence and repeat penalties disabled. Its Hermes-derived agent behavior and large context make it a candidate for code explanation, multi-file planning, and structured technical work. The README does not provide benchmark results, so test it against your own coding tasks before relying on it for production code.

Tool-oriented assistant workflows. The Hermes agent transfer and the recommended function-calling dataset commands make this release relevant to local agents that need to follow tool-use conventions. The model card also gives a JSON-oriented system prompt pattern for schema-constrained responses. Treat these as prompting guidance, not proof of reliable tool execution: the README reports no function-calling accuracy or structured-output benchmark.

Long-document analysis. The native context is 262K tokens, and the model card says to keep at least 128K context to preserve thinking capabilities. That makes it suitable for experiments involving long reports, codebases, or document collections, subject to runtime memory and context-cache costs. The model card says YaRN can extend context to 1M, but it provides no quality or speed measurements at that length.

Uncensored creative writing and role-play. The base is explicitly described as uncensored, and the maintainer reports 0 refusals on a 465-prompt test for the base model. That result is a maintainer-reported test, not an independent safety or quality evaluation. The card supplies a non-thinking creative profile—temperature 0.7, top_p 0.8, top_k 20, min_p 0.01—and optional creative system prompts.

Multimodal local experiments. The architecture is described as natively multimodal for text, image, and video. For vision in GGUF, the README requires an mmproj file alongside the main model file. The provided material does not specify supported image or video formats, resolution limits, or multimodal benchmark results, so validate those capabilities in the chosen runtime before building a workflow around them.

Limitations

The README gives no independent benchmark scores, measured inference speed, latency, throughput, or exact VRAM requirements. The 3B active-parameter figure does not mean the full 35B model fits in 3 GB of memory; quantization, context length, cache precision, and CPU/GPU offload all affect resource use.

The maintainer recommends F16 K and V cache quantization for APEX, maximum GPU offload, and forcing MoE weights for 40 layers onto CPU, but does not provide a hardware configuration or performance result for those settings.

The model card’s Genesis claims—including noise reduction, preservation of 99% of signal and learned gradient, and improved stability—are the maintainer’s description of a custom tensor-repair method. The supplied material does not include controlled comparisons, reproducible evaluation results, or independent validation. The model is not described as retrained or conventionally fine-tuned end-to-end: the maintainer says Genesis modifies GGUF tensor data and that Hermes data was transferred from around 2,000 blocks in two FFN expert tensors.

The uncensored base and the reported 0/465 refusal result signal a safety tradeoff. Do not assume the model will refuse harmful requests or meet a deployment’s safety requirements. The README also warns that context below 128K may affect thinking capabilities, while the 1M YaRN extension has no reported quality measurements. Vision requires a separate mmproj file; the main GGUF alone is not enough for that use.

The listed license is Apache-2.0. The provided material does not clarify whether every upstream component and redistributed quantization carries identical terms, so review the relevant model and artifact licenses before commercial use. No specific newer model is identified as having superseded this release.

How It Compares

The supplied alternatives are other Genesis Hermes versions from the same maintainer. The provided information does not give controlled quality, speed, or cost comparisons between versions, so version numbers alone cannot establish which performs better.

  • Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF: Pick this release when you want the specific combination documented here: the HauhauCS uncensored base, Hermes data transfer from about 2,000 blocks in two FFN expert tensors, and the described Genesis repair. Pick V6 instead if its own model card documents a feature, quant, or runtime fit that your deployment needs; the supplied comparison data does not establish a quality or speed winner.
  • Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-Final-GGUF: Choose this model when you need the exact release described in this guide and its stated settings and quant recommendations. Choose Final when its artifact or card better matches your runtime or deployment requirements; no measured quality, latency, or memory comparison is provided here.
  • Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V10-GGUF: This release is the documented choice if you want the current model’s stated 262K native context, multimodal support, and APEX/NVFP4 guidance. Consider V10 if its own card or files provide a better fit for your task; the supplied material does not establish that V10 is faster, more accurate, or cheaper to run.
  • Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V9-GGUF: Pick this release for the specific Genesis/Hermes construction and runtime instructions documented here. Pick V9 if you have validated it on your workload or need a version-specific artifact; no benchmark evidence in the supplied material supports a general quality or speed preference.
  • Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF: This release has explicit recommendations for APEX settings, context length, and multimodal mmproj use. V7 may suit you if its own files or card match your deployment better, but the provided information does not quantify differences in output quality, inference speed, or cost.

Technical Specifications

The model uses a mixture-of-experts architecture with 35B total parameters and about 3B active per forward pass. It has 256 experts, with 8 routed experts and 1 shared expert per token. Its 40 layers follow a repeating pattern of three DeltaNet-MoE layers followed by one Attention-MoE layer: 10 repetitions in total. The attention design combines Gated DeltaNet linear attention and full softmax attention in a 3:1 ratio.

  • Context: 262K tokens natively; extendable to 1M with YaRN. The model card recommends at least 128K context to preserve thinking capabilities.
  • Modalities: Text, image, and video are listed as native capabilities. Vision support in GGUF requires an mmproj file alongside the main GGUF.
  • Vocabulary and languages: 248K vocabulary; 201 languages.
  • Model format and runtime: GGUF; compatible with llama.cpp, LM Studio, koboldcpp, and other GGUF-compatible runtimes. Use llama.cpp’s --jinja flag for chat-template handling.
  • Quantization: Recommended options are NVFP4 and APEX. APEX Compact is recommended for GPUs with 8 GB or 12 GB of memory. The README links a quantization script with Unsloth profile support and points to a separate NVFP4 GGUF release.
  • APEX settings: Set K and V cache quantization to F16, force MoE weights for 40 layers onto CPU, set GPU offload to maximum, and set active experts to 8.
  • Sampling profiles: Thinking-mode Hermes agent and coding/precise tasks use temperature 0.6, top_p 0.95, top_k 20, min_p 0.01, seed 42, and disabled presence penalty; the Hermes agent profile uses repeat_penalty 1.05, while coding/precise tasks disable repeat penalty. Thinking-mode general uses temperature 0.95 with the same top_p, top_k, min_p, seed, and disabled penalties. Non-thinking creative general uses temperature 0.7, top_p 0.8, top_k 20, min_p 0.01, seed 42, and disabled penalties.
  • Genesis process, as described by the maintainer: The maintainer says the process scans ssm_conv1d tensors and adjusts head balance; scans tensor blocks in chunks using three parameters and replaces zero chunks with a selected chunk; then detects and reduces training noise with custom SVD using the Marchenko–Pastur law. The process excludes token_embd.weight, output.weight, 1D tensors, biases, and norms. The maintainer says it runs on GGUF using Python and a free Google Colab Tesla T4 GPU, without retraining or fine-tuning the full model.
  • Training and data: The README identifies the HauhauCS uncensored model as the base and a Hermes Qwen3.5 GGUF fine-tune as the source of transferred data. It reports around 2,000 blocks from two FFN expert tensors. It gives no training-step count, full dataset size, or training compute for the base model.
  • Performance and memory: No measured tokens per second, latency, batch-size guidance, or exact VRAM requirement is provided. The 8 GB and 12 GB figures refer to the maintainer’s recommendation to use APEX Compact, not a guarantee that every context length or runtime configuration will fit.

Model Inputs and Outputs

Inputs

  • Text: Prompts and conversation messages using the supplied chat template. In llama.cpp, enable --jinja for proper template handling.
  • Images: Supported as a native modality according to the model card; GGUF vision use requires the matching mmproj file. The README does not specify image formats or resolution limits.
  • Video: Listed as a native modality. The README does not specify video container formats, frame handling, or limits.
  • Context: Native context is 262K tokens. YaRN extension to 1M is stated, but the card recommends at least 128K to preserve thinking capabilities.

Outputs

  • Text: Generated assistant responses, including coding, general, creative, and agent-style responses.
  • Structured text: The model card recommends a JSON system prompt with a supplied schema placeholder for agentic tasks. It does not report a structured-output guarantee or validation rate.
  • Multimodal responses: The model is described as natively multimodal, but the README does not specify output modalities or post-processing requirements beyond the mmproj requirement for vision input.

Getting Started

The README provides runtime and sampling settings, but no Python loading or inference example, and it does not identify a Python inference library or API. A reliable copy-pasteable Python example cannot be derived from the supplied information. For a documented local setup, use a GGUF-compatible runtime such as llama.cpp, LM Studio, or koboldcpp; in llama.cpp, pass --jinja, and include the mmproj file when using vision. Start with the recommended system prompt and keep context at or above 128K if you want to follow the maintainer’s guidance for thinking mode.

Frequently Asked Questions

Q: Can I use Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-GGUF commercially?

A: The listed license is Apache-2.0. The supplied material does not clarify the licensing of every upstream component or quantized artifact, so review those terms before commercial deployment.

Q: What hardware or VRAM do I need to run it?

A: The README gives no exact VRAM requirement. It recommends APEX Compact for 8 GB and 12 GB GPUs, and gives APEX offload settings, but those recommendations do not guarantee that every context length and runtime configuration will fit.

Q: Is this model suitable for local tool-calling agents?

A: It includes Hermes-agent behavior transferred from a Hermes fine-tune, and the model card recommends a JSON system prompt and Hermes function-calling commands. It provides no tool-calling benchmark, so validate tool selection and argument formatting for your application.

Q: What are the known failure modes or quality issues?

A: The README does not report measured failure rates or benchmark results. It describes Genesis as a tensor-repair method intended to reduce noise and improve stability, but those claims are not accompanied by independent evaluation; the uncensored base also means you should not rely on it to refuse unsafe requests.

Q: Can I fine-tune this model, and which framework supports it?

A: The README does not document a fine-tuning workflow or name a fine-tuning framework. It links a quantization script with Unsloth profile support, which is for quantization rather than evidence of a supported fine-tuning path.

Q: What input format does the model expect?

A: It uses GGUF-compatible runtimes and a chat template; the model card recommends --jinja with llama.cpp. For vision, place the mmproj file alongside the main GGUF. The README does not specify image or video file formats.

Q: How fast is inference, and what batch sizes are practical?

A: The supplied material reports no inference speed, latency, or batch-size measurements. Performance will depend on the quantization, runtime, hardware, context length, and offload settings.

Q: Does it support long context and thinking mode?

A: The stated native context is 262K tokens, with extension to 1M using YaRN. The maintainer recommends at least 128K context to preserve thinking capabilities, but provides no quality measurements for the 1M extension.

Feature Image source: AIModels.fyi

This is a simplified guide to an AI model called Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-GGUF maintained by LuffyTheFox. If you like these kinds of analysis, join AIModels.fyi or follow us on Twitter.



Source link