Charset vs Collation DSN Parameters in go-sql-driver/mysql: Impact Explained

In the Go MySQL driver, the charset DSN parameter defines the connection's character encoding while collation specifies the sorting and comparison rules, determining whether the driver issues SET NAMES <charset> or SET NAMES <charset> COLLATE <collation> during connection initialization.

The go-sql-driver/mysql package parses Data Source Name (DSN) strings into a Config struct to establish how your application communicates character encoding to the MySQL server. Understanding the distinct roles of these two parameters prevents "Incorrect string value" errors and ensures consistent sorting behavior across queries.

How the Driver Parses charset and collation

The distinction begins in dsn.go, where the parseDSNParams function handles DSN string parsing differently for each parameter.

The Config Structure

In dsn.go, the driver maintains separate fields for these values:

  • cfg.charsets (slice): Populated by splitting the charset value on commas using strings.Split(value, ",")【/dsn.go#L44-L47】. This allows specifying multiple fallback character sets.
  • cfg.Collation (string): Assigned directly as a scalar value in parseDSNParams【/dsn.go#L49-L52】.

DSN Parameter Processing

When parsing the DSN, the driver treats charset as a potentially multi-value parameter while collation accepts only a single value:

// From dsn.go - charset accepts comma-separated values
cfg.charsets = strings.Split(value, ",")

// From dsn.go - collation accepts single value
cfg.Collation = value

SQL Commands Generated During Connection

Upon establishing a connection, the driver executes character set initialization commands based on which parameters you provide.

Only charset specified:

SET NAMES utf8mb4

Both charset and collation specified:

SET NAMES utf8mb4 COLLATE utf8mb4_unicode_ci

If collation is omitted, the driver does not append a COLLATE clause, allowing MySQL to use the server's default collation for the specified character set【/dsn.go#L13-L17】.

Impact on Application Behavior

The parameters affect different aspects of database operations:

Character Encoding (charset)

  • Controls byte representation: Determines how string literals and result sets are encoded (e.g., utf8mb4, latin1). Incorrect settings cause data corruption or "Incorrect string value" errors when inserting multibyte characters.
  • Affects prepared statements: Governs how the driver encodes query parameters before transmission.

Collation Rules (collation)

  • Controls comparison semantics: Determines case-sensitivity, accent sensitivity, and language-specific sorting rules (e.g., utf8mb4_unicode_ci vs utf8mb4_general_ci).
  • Affects WHERE clauses and ORDER BY: Changes which rows match string comparisons and how results are sorted without changing the underlying data encoding.
  • Performance considerations: Unicode collations like utf8mb4_unicode_520_ci perform more complex computations than utf8mb4_general_ci, potentially impacting query speed for large text comparisons.

Configuration Examples

Explicitly set both parameters for consistent cross-platform behavior:

package main

import (
	"database/sql"
	"log"

	_ "github.com/go-sql-driver/mysql"
)

func main() {
	// Full Unicode support with specific collation rules
	dsn1 := "root@tcp(127.0.0.1:3306)/test?charset=utf8mb4&collation=utf8mb4_unicode_ci"
	db1, err := sql.Open("mysql", dsn1)
	if err != nil {
		log.Fatal(err)
	}
	// Issues: SET NAMES utf8mb4 COLLATE utf8mb4_unicode_ci

	// Server default collation (often utf8mb4_general_ci)
	dsn2 := "root@tcp(127.0.0.1:3306)/test?charset=utf8mb4"
	db2, err := sql.Open("mysql", dsn2)
	if err != nil {
		log.Fatal(err)
	}
	// Issues: SET NAMES utf8mb4
}

Verify the generated commands by enabling MySQL's general query log—look for the SET NAMES statements immediately following connection establishment.

Key Source Files

The implementation spans three critical files in the go-sql-driver/mysql repository:

  • dsn.go: Defines the Config struct with charsets and Collation fields, and implements parseDSNParams parameter handling.
  • collations.go: Contains the unsafeCollations map used when validating collation safety with interpolateParams.
  • driver.go: Executes the SET NAMES command during connection initialization based on the parsed configuration.

Summary

  • charset ([]string in dsn.go) sets the byte encoding; the driver may accept multiple values but uses the first for SET NAMES.
  • collation (string in dsn.go) sets comparison rules; only used when explicitly provided, triggering a COLLATE clause in the initialization command.
  • Impact: Charset prevents encoding errors during data insertion, while collation controls case-sensitivity and sorting in WHERE clauses and ORDER BY operations.
  • Omission: Leaving collation empty causes MySQL to apply its default collation for the specified charset, which varies by server configuration.

Frequently Asked Questions

Can I specify multiple character sets in the DSN?

Yes. The parseDSNParams function in dsn.go splits the charset value on commas, storing it as a string slice in cfg.charsets. However, the connection initialization typically uses the first charset in the slice for the SET NAMES command, with additional charsets potentially serving as fallbacks depending on client implementation details.

What happens if I set collation without specifying charset?

The SET NAMES SQL command requires a character set name. While the driver stores the collation value separately in cfg.Collation, MySQL cannot apply a collation without knowing which character set it belongs to. Specify at least one charset (e.g., utf8mb4) to ensure proper connection initialization.

Which collation should I use with utf8mb4?

Use utf8mb4_unicode_ci for standard Unicode sorting rules that handle international characters correctly, or utf8mb4_0900_ai_ci for MySQL 8.0+ improved accuracy. Avoid utf8mb4_general_ci for applications requiring accurate language-specific sorting, though it offers marginally better performance.

Does changing the collation DSN parameter affect database performance?

Yes. Case-insensitive Unicode collations (*_unicode_ci, *_0900_ai_ci) require more CPU resources for string comparison than binary (*_bin) or general (*_general_ci) collations. If your application performs heavy text filtering or sorting on large datasets, choose simpler collations like utf8mb4_general_ci, or use binary collation for exact matching and handle normalization in application logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →