Known Issues with Qwen3 Coder Next on Unified Memory in ODS
ODS automatically routes away from the Qwen3 Coder Next model on unified-memory hosts to prevent a critical correctness failure where the model generates only invalid "?" tokens, substituting the Qwen3.6-35B-A3B-UD model instead according to repository policy.
The Osmantic/ODS repository treats the qwen3-coder-next model as incompatible with unified system memory architectures found in APU-style AMD processors, NVIDIA unified-memory GPUs, and Apple Silicon devices. When deployed on these hardware configurations, the model exhibits a severe token generation pathology that renders it unusable for production inference tasks.
Root Cause of the Correctness Pathology
On hosts exposing unified memory envelopes, Qwen3 Coder Next produces entirely invalid output consisting solely of "?" tokens rather than coherent code or text generation. This hardware-specific failure mode affects APU-style AMD architectures, NVIDIA unified-memory tiers, and Apple Silicon environments where system and GPU memory share a unified address space.
The issue manifests as a total loss of model capability rather than degraded performance, making automatic detection and substitution critical for reliable deployment.
Policy-Driven Model Substitution Mechanism
ODS implements a hardcoded policy that intercepts model selection for unified-memory hardware and redirects requests to a compatible alternative. The repository maintains explicit substitution logic across multiple implementation layers to ensure consistent behavior.
Python Model Selector Implementation
The primary selection logic resides in ods/scripts/select-model.py, where the UNIFIED_MEMORY_POLICY constant defines the substitution rule. When the selector detects a unified-memory host, it swaps the coder-next entry for the larger Qwen3.6-35B-A3B-UD model to avoid the failure.
# ods/scripts/select-model.py
# Unified‑memory hosts (Strix Halo SH_LARGE, future AMD/NV unified‑memory tiers)
# hit the same coder‑next correctness pathology as Spark aarch64.
# Until upstream fixes coder‑next on unified‑memory backends, route the
# qwen profile to the same 35B‑A3B substitution used for Spark — same
# model id, separate policy tag so the recommendation_reason is honest
# about why the substitution fired.
UNIFIED_MEMORY_POLICY = "unified‑memory‑coder‑next‑a3b‑v1"
UNIFIED_MEMORY_MODEL_ID = SPARK_AARCH64_MODEL_ID
This policy tag ensures the recommendation_reason field transparently documents why the substitution occurred, maintaining auditability for deployment decisions.
Bash Tier-Map Enforcement
The Bash implementation in ods/installers/lib/tier-map.sh mirrors this policy for installer scripts and command-line tooling. The library checks hardware detection results and forces the A3B model selection when unified memory is present.
# ods/installers/lib/tier-map.sh
if [[ "$memory_type" == "unified" && "$gpu_type" == "amd" ]]; then
LLM_MODEL="qwen3.6-35b-a3b"
GGUF_FILE="qwen3.6-35b-a3b-ud-q4.gguf"
fi
The substitution applies specifically to tiers like SH_LARGE (Strix Halo large configuration), ensuring users with high-end APU hardware receive a functional model rather than the compromised coder-next variant.
Contract Testing and Verification
ODS validates the unified-memory substitution through contract tests that enforce the policy at the CI/CD level. The Windows-specific tier-map contract in ods/tests/contracts/test-windows-strix-tier-map.sh verifies that the PowerShell resolver selects the A3B model when processing unified-memory configurations.
# ods/tests/contracts/test-windows-strix-tier-map.sh
if grep -q '\$script:UNIFIED_MEMORY_POLICY = "unified‑memory‑coder‑next‑a3b‑v1"' "$TIER_MAP" \
&& grep -q 'RESOLVED=qwen3.6-35b-a3b|Qwen3.6-35B-A3B‑UD‑Q4_K_M.gguf|context‑aware‑largest‑capable‑general‑v1+unified‑memory‑coder‑next‑a3b‑v1' <<<"$OUT"; then
pass "PowerShell resolver selects Qwen3.6 A3B with unified‑memory policy"
These tests guarantee that modifications to the tier-map logic cannot inadvertently expose unified-memory hosts to the coder-next pathology, preventing regression in production deployments.
Impact on Hardware Configurations
The routing policy affects specific high-end hardware configurations including:
- AMD Strix Halo (SH_LARGE tier): APU-style unified memory architectures
- Future AMD unified-memory tiers: Upcoming implementations sharing system and video memory
- NVIDIA unified-memory GPUs: Devices exposing unified memory envelopes
- Apple Silicon: macOS hosts with unified memory architecture
According to the repository documentation in ods/README.md, this exclusion is deliberate and documented: "Unified‑memory hosts are routed away from qwen3‑coder‑next when that model would otherwise be selected, because current repo policy documents correctness issues on those backends."
Summary
- ODS substitutes Qwen3 Coder Next with Qwen3.6-35B-A3B-UD on all unified-memory hosts to prevent invalid token generation
- The correctness pathology causes the model to emit only "?" tokens on APU-style AMD, NVIDIA unified-memory, and Apple Silicon hardware
- Policy enforcement occurs in
select-model.pyviaUNIFIED_MEMORY_POLICYandUNIFIED_MEMORY_MODEL_IDconstants - Bash tier-map logic in
tier-map.shimplements hardware detection and automatic model substitution for installer scripts - Contract tests in
test-windows-strix-tier-map.shvalidate that the resolver never selects coder-next for unified-memory configurations
Frequently Asked Questions
Why does Qwen3 Coder Next fail on unified memory systems?
Qwen3 Coder Next exhibits a hardware-specific correctness pathology on unified-memory architectures where the model produces invalid output consisting entirely of "?" tokens rather than coherent generation. This affects APU-style AMD processors, NVIDIA unified-memory GPUs, and Apple Silicon devices where CPU and GPU share memory address space.
What model does ODS use instead of Qwen3 Coder Next on unified memory hosts?
ODS substitutes the Qwen3.6-35B-A3B-UD model (also referred to as qwen3.6-35b-a3b-ud-q4) when detecting unified memory hardware. This substitution occurs automatically through the UNIFIED_MEMORY_POLICY mechanism defined in ods/scripts/select-model.py and enforced in ods/installers/lib/tier-map.sh.
How can I verify if my system is using the unified memory substitution policy?
Check the model recommendation output for the policy tag unified‑memory‑coder‑next‑a3b‑v1 appended to the recommendation reason. When running ods select-model on a unified-memory host, the output will show SELECTED_MODEL=qwen3.6-35b-a3b-ud-q4 and include the unified-memory policy identifier in the recommendation metadata.
Will ODS remove this restriction once the upstream issue is fixed?
The source code comments in ods/scripts/select-model.py indicate this is a temporary workaround "until upstream fixes coder-next on unified-memory backends." However, the repository maintains contract tests and explicit policy constants to ensure the substitution remains active until an explicit update removes the UNIFIED_MEMORY_POLICY enforcement.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →