How Nitter Parses the Legacy Twitter API Tweet Format

Nitter parses the legacy Twitter API tweet format by inspecting JSON payloads for a "legacy" key in src/parser.nim, routing the nested object through parseTweet to extract core fields, and then calling parseLegacyMediaEntities to resolve attached photos, videos, and GIFs.

Nitter is an open-source alternative Twitter front-end that consumes data from Twitter's unofficial endpoints. Understanding how Nitter parses the legacy Twitter API tweet format is essential for anyone auditing its data pipeline or contributing to the project. All parsing logic is centralized in src/parser.nim, where raw JSON is transformed into strongly typed Nim objects defined in src/types.nim.

Detecting the Legacy Twitter API Tweet Format in parseGraphTweet

Before any field extraction occurs, Nitter determines whether an incoming JSON node represents a legacy tweet. This happens inside the parseGraphTweet procedure in src/parser.nim. The function checks for the presence of a "legacy" field or a GraphQL "rest_id" fallback. If neither exists, it returns an empty Tweet() immediately.

if "legacy" notin js and "rest_id" notin js:
  return Tweet()
var jsCard = select(js{"card"}, js{"tweet_card"}, js{"legacy", "tweet_card"})
result = parseTweet(js{"legacy"}, jsCard, replyId, hasArticle)

When the "legacy" key is present, the inner object is passed to parseTweet along with the extracted card data. This branching logic ensures that legacy payloads are handled separately from newer GraphQL structures while still sharing the same internal Tweet type. The detection gate is located around lines 85‑90 of src/parser.nim.

Mapping Legacy Twitter API Tweet Fields in parseTweet

The parseTweet procedure in src/parser.nim performs the bulk of the field mapping. It accepts the legacy JSON node and constructs a fully populated Tweet object by reading classic attributes such as id_str, full_text, and created_at. It also resolves reply metadata and populates a TweetStats instance with engagement counters.

Key mappings implemented in parseTweet include:

  • Tweet ID: extracted from js{"id_str"}.getId
  • Text content: read from js{"full_text"}.getStr
  • Timestamp: parsed from created_at or created_at_ms
  • Statistics: reply_count, retweet_count, and favorite_count mapped to TweetStats
proc parseTweet(js: JsonNode; ...): Tweet =
  result = Tweet(
    id: js{"id_str"}.getId,
    text: js{"full_text"}.getStr,
    time: if js{"created_at"}.notNull:
            js{"created_at"}.getTime
          else:
            js{"created_at_ms"}.getTimeFromMs,
    stats: TweetStats(
      replies: js{"reply_count"}.getInt,
      retweets: js{"retweet_count"}.getInt,
      likes: js{"favorite_count"}.getInt
    )
  )

This procedure spans roughly lines 81‑112 in src/parser.nim and serves as the primary bridge between Twitter's legacy schema and Nitter's internal representation.

Resolving Media from the Legacy Twitter API Tweet Format

After the core Tweet struct is assembled, Nitter processes media attachments by calling parseLegacyMediaEntities on the original JSON node. This procedure walks the extended_entities.media array and creates a Media variant for each entry. The variant distinguishes between photo, video, and animated_gif types based on the type field in the legacy payload.

parseLegacyMediaEntities(js, result)

Inside the procedure, the code iterates over the media array as shown below:

with jsMedia, js{"extended_entities", "media"}:
  for m in jsMedia:
    case m.getTypeName
    of "photo": # create Photo media

    of "video": # create Video media with attribution

    of "animated_gif": # create Gif media

The resulting Media objects are appended directly to Tweet.media, making them available for timeline and detail rendering. Video entries also carry attribution data when present in the legacy entity. This logic is implemented in lines 107‑133 of src/parser.nim.

End-to-End Legacy Tweet Parsing Example

In practice, the entire flow from raw API response to structured tweet is orchestrated within src/parser.nim. A resolver fetches JSON from an endpoint such as https://api.twitter.com/2/timeline/... and passes it to parseGraphTweet. The function auto-detects the legacy wrapper, strips it away, and delegates to parseTweet and parseLegacyMediaEntities in sequence.

import packedjson, parser

let payload = fetchTwitterAPI("/2/timeline/...")   # returns JsonNode

let tweet = parseGraphTweet(payload)               # auto-detects legacy

echo tweet.text                                    # full text

echo tweet.stats.likes                             # like count

echo tweet.media.len                               # attached media count

You can also extract and inspect media variants after parsing:

let tweet = parseTweet(js{"legacy"})
for m in tweet.media:
  case m.kind
  of photoMedia: echo "Photo:", m.photo.url
  of videoMedia: echo "Video:", m.video.url
  of gifMedia:   echo "GIF:",   m.gif.url

Helper utilities from src/parserutils.nim—including getStr, getId, and select—power the safe JSON navigation throughout these procedures. For GraphQL-specific pathways that occasionally fall back to legacy parsing, see src/experimental/parser/graphql.nim.

Summary

  • Nitter detects legacy payloads in src/parser.nim by checking for a top-level "legacy" key inside parseGraphTweet.
  • The parseTweet procedure maps legacy fields such as id_str, full_text, and favorite_count to the internal Tweet and TweetStats types.
  • Media attachments are resolved by parseLegacyMediaEntities, which iterates over extended_entities.media and produces photo, video, or animated_gif variants.
  • The entire pipeline is self-contained in src/parser.nim and relies on helpers from src/parserutils.nim and type definitions from src/types.nim.

Frequently Asked Questions

What file handles the legacy Twitter API tweet parsing in Nitter?

The central parsing logic lives in src/parser.nim. This file contains the parseGraphTweet, parseTweet, and parseLegacyMediaEntities procedures that detect, map, and enrich legacy tweet payloads. Supporting type definitions are located in src/types.nim, and JSON navigation helpers are provided by src/parserutils.nim.

How does Nitter distinguish between legacy and GraphQL tweet payloads?

Inside parseGraphTweet, Nitter checks whether the JSON node contains a "legacy" key or a GraphQL "rest_id" key. If neither is present, it returns an empty Tweet(). When "legacy" is found, the nested object is sent to the legacy parser; otherwise, the GraphQL pathway in src/experimental/parser/graphql.nim takes over.

Which legacy JSON fields does Nitter map to the internal Tweet type?

According to the source code in src/parser.nim, Nitter maps id_str to the tweet ID, full_text to the text content, created_at or created_at_ms to the timestamp, and reply_count, retweet_count, and favorite_count to the TweetStats counters. These mappings occur inside the parseTweet procedure.

How does Nitter handle media attachments from legacy tweets?

Nitter calls parseLegacyMediaEntities after building the core tweet object. This procedure reads the extended_entities.media array from the legacy JSON and generates a Media variant for each attachment. It supports photos, videos, and animated GIFs, and also extracts video attribution when available.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →