How to Troubleshoot Common Errors and Connection Issues in MLX Omni Server

Use the RequestResponseLoggingMiddleware logs and verify your MLX_OMNI_CORS environment variable to resolve most startup, connection, and HTTP 500 failures.

MLX Omni Server is a FastAPI-based inference gateway that exposes chat, embedding, image, and audio endpoints. When you need to troubleshoot common errors and connection issues in MLX Omni Server, the solution typically lies in the middleware stack, the Uvicorn binding configuration, or the environment-specific CORS settings.

Diagnostic Categories and Root Causes

Most runtime failures fall into three architectural layers: the network binding layer (Uvicorn/server startup), the HTTP middleware layer (CORS and logging), and the application service layer (model adapters and file I/O). Understanding which layer emits the error lets you jump directly to the relevant source file.

Server Startup and Port Binding Errors

If the process exits immediately with OSError: [Errno 48] Address already in use, the port specified in main.py is occupied. In src/mlx_omni_server/main.py, the server invokes uvicorn.run() with default arguments binding to 0.0.0.0:10240.


# Check what is using port 10240

lsof -i :10240

# Start on a free port

python -m mlx_omni_server.main --port 10241

Permission errors binding to privileged ports (e.g., 80) also surface here; either use a higher port or run with appropriate privileges.

CORS and Browser Connectivity Failures

Browser consoles showing "CORS policy: No 'Access-Control-Allow-Origin' header" indicate the CORSMiddleware is misconfigured. The server constructs this middleware in configure_cors_middleware() inside src/mlx_omni_server/main.py based on the MLX_OMNI_CORS environment variable or the --cors-allow-origins CLI flag.


# Allow all origins (development only)

export MLX_OMNI_CORS="*"
python -m mlx_omni_server.main --cors-allow-origins="*"

If the variable is unset, the middleware defaults may block your frontend domain. Always verify the middleware is reloaded after changing environment variables.

HTTP 500 Errors and Unhandled Exceptions

Internal server errors originate in the service layer but propagate through the middleware stack. The RequestResponseLoggingMiddleware in src/mlx_omni_server/middleware/logging.py catches unhandled exceptions at line 14-15, logs the full traceback, and re-raises it.


# Excerpt from middleware/logging.py showing the exception trap

try:
    response = await call_next(request)
except Exception as e:
    logger.exception("Unhandled exception: %s", e)
    raise

When you see a 500 response, check the logs for the ERROR line immediately preceding the request summary. Common triggers include missing API keys in chat/openai/openai_adapter.py (lines 377 and 499 wrap model calls in exception handlers) or invalid model identifiers in src/mlx_omni_server/embeddings/router.py.

Connection Refused and Timeout Issues

curl: (7) Failed to connect means the client cannot reach the TCP socket. Verify the host binding in main.py: the default 0.0.0.0 accepts external connections, but 127.0.0.1 restricts traffic to localhost. Firewalls and Docker network isolation also cause this symptom.


# Test local binding only

curl http://127.0.0.1:10240/health

# Verify the server is listening on all interfaces

netstat -tlnp | grep 10240

Timeouts during long inference requests (e.g., image generation in src/mlx_omni_server/images/images_service.py) may require increasing the Uvicorn --timeout-keep-alive value or the client-side timeout.

Step-by-Step Troubleshooting Workflow

Follow this sequence to isolate the failure point:

  1. Validate the process is running

    ps aux | grep -E "uvicorn|mlx_omni"
  2. Inspect real-time logs The middleware logs every request. Set MLX_OMNI_LOG_LEVEL=debug to capture detailed timing:

    export MLX_OMNI_LOG_LEVEL=debug
    python -m mlx_omni_server.main
  3. Check the health endpoint

    curl -i http://localhost:10240/health

    A 200 OK confirms the router stack in src/mlx_omni_server/routers.py is mounted correctly. A 404 indicates a routing registration failure.

  4. Force a controlled error Send a request with an invalid model name to verify error propagation:

    curl -X POST http://localhost:10240/v1/embeddings \
      -H "Content-Type: application/json" \
      -d '{"model": "nonexistent-model", "input": "test"}'

    Expect a logged exception pointing to embeddings/router.py or the underlying service.

  5. Review CORS headers manually

    curl -I -H "Origin: http://example.com" http://localhost:10240/health

    Look for access-control-allow-origin in the response headers. Absence confirms a CORS configuration gap in main.py.

Key Source Files for Debugging

File Purpose Link
src/mlx_omni_server/main.py Entry point, Uvicorn startup, CORS middleware setup main.py
src/mlx_omni_server/routers.py Aggregates all sub-routers (chat, embeddings, images, TTS, STT) routers.py
src/mlx_omni_server/middleware/logging.py Captures unhandled exceptions and request metrics logging.py
src/mlx_omni_server/chat/openai/openai_adapter.py Wraps OpenAI SDK calls with exception handling openai_adapter.py
src/mlx_omni_server/embeddings/router.py Embedding endpoint definitions and validation embeddings/router.py
src/mlx_omni_server/images/images_service.py Image generation logic; raises model-specific exceptions images_service.py

Summary

  • Port conflicts surface in main.py where uvicorn.run() binds to 10240 by default; use --port to override.
  • CORS errors require setting MLX_OMNI_CORS or --cors-allow-origins before the CORSMiddleware initializes in main.py.
  • HTTP 500 responses are logged by RequestResponseLoggingMiddleware in middleware/logging.py; trace the exception back to the specific service adapter.
  • Connection refused indicates network or binding misconfiguration; verify the host argument and firewall rules.
  • Model-specific failures (e.g., embeddings, images) propagate from service files like images_service.py and are caught by the global exception handlers.

Frequently Asked Questions

Why does my browser show a CORS error even though the server is running?

The CORSMiddleware in src/mlx_omni_server/main.py requires explicit origin whitelisting via the MLX_OMNI_CORS environment variable or the --cors-allow-origins flag. Without this, the server omits Access-Control-Allow-Origin headers, causing browsers to block cross-origin requests.

How do I find the root cause of a 500 Internal Server Error?

Check the logs captured by RequestResponseLoggingMiddleware in src/mlx_omni_server/middleware/logging.py (lines 14-15). It logs the full traceback before re-raising the exception, allowing you to identify whether the failure occurred in a model adapter like openai_adapter.py or in a downstream file operation.

What should I do if the server fails to start with "Address already in use"?

The default configuration in main.py attempts to bind port 10240. Identify the conflicting process with lsof -i :10240, terminate it, or start the server on an alternative port using python -m mlx_omni_server.main --port 10241.

Why are my requests timing out during image generation?

Long-running inference in src/mlx_omni_server/images/images_service.py can exceed default HTTP timeouts. Increase the Uvicorn keep-alive setting or implement client-side retry logic, and ensure your network infrastructure does not impose stricter timeouts than the server configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →