go go/book/learninggo

Strings

In Go, a string is a sequence of bytes (not characters), and it represents text data. Strings are immutable, meaning once they are created, their content cannot be modified.

  • Representation: A string in Go is represented as a sequence of UTF-8 encoded bytes.
  • Immutable: You cannot change individual characters in a string directly.
  • Length: The len() function returns the number of bytes, not the number of characters.
  • UTF-8 Encoding: Strings are encoded in UTF-8 by default, which supports Unicode but may require multiple bytes per character.
package main
 
import "fmt"
 
func main() {
    s := "Hello, 世界"   // A string containing ASCII and Unicode characters
    fmt.Println(s)      // Output: Hello, 世界
    fmt.Println(len(s)) // Output: 13 (in bytes, not characters)
}

In this example, “世界” (meaning “world” in Chinese) requires more than one byte per character in UTF-8, hence the length 13 in bytes.

Runes

A rune is an alias for int32 in Go and is used to represent a single Unicode code point. Rune is useful when you need to work with individual characters in a string, especially those outside the ASCII range.

  • Representation: rune represents a single Unicode code point as a 32-bit integer.
  • Character Encoding: Each rune represents a character (regardless of how many bytes it takes).
  • Mutable Equivalent: While strings are immutable, you can use []rune to create a mutable slice of characters from a string.
  • UTF-8 Decoding: Converting a string to []rune decodes each character, handling multi-byte characters correctly.
package main
 
import "fmt"
 
func main() {
    s := "Hello, 世界"
    
    // Convert string to slice of runes
    runes := []rune(s)
 
    // Print each rune as a character and Unicode code point
    for i, r := range runes {
        fmt.Printf("Character %d: %c, Unicode: %U\n", i, r, r)
    }
}

Key Differences

Featurestringrune
TypeImmutable sequence of bytesAlias for int32, represents a single code point
EncodingUTF-8Represents one Unicode character
LengthReturns number of bytes (not characters)Each rune is a single code point
UsageFor text storage and manipulationFor handling single characters
Conversion[]rune(string) to decode charactersstring([]rune) to convert back to a string

When to Use Each

  • string: For most text handling, storing, and displaying, as it’s efficient and suited for UTF-8 encoded data.
  • rune: When you need to handle individual characters, work with Unicode, or manipulate text that may contain multi-byte characters.