How Nitter Parses the Legacy Twitter API Tweet Format
Nitter parses the legacy Twitter API tweet format by inspecting JSON payloads for a "legacy" key in src/parser.nim, routing the nested object through parseTweet to extract core fields, and then calling parseLegacyMediaEntities to resolve attached photos, videos, and GIFs.
Nitter is an open-source alternative Twitter front-end that consumes data from Twitter's unofficial endpoints. Understanding how Nitter parses the legacy Twitter API tweet format is essential for anyone auditing its data pipeline or contributing to the project. All parsing logic is centralized in src/parser.nim, where raw JSON is transformed into strongly typed Nim objects defined in src/types.nim.
Detecting the Legacy Twitter API Tweet Format in parseGraphTweet
Before any field extraction occurs, Nitter determines whether an incoming JSON node represents a legacy tweet. This happens inside the parseGraphTweet procedure in src/parser.nim. The function checks for the presence of a "legacy" field or a GraphQL "rest_id" fallback. If neither exists, it returns an empty Tweet() immediately.
if "legacy" notin js and "rest_id" notin js:
return Tweet()
var jsCard = select(js{"card"}, js{"tweet_card"}, js{"legacy", "tweet_card"})
result = parseTweet(js{"legacy"}, jsCard, replyId, hasArticle)
When the "legacy" key is present, the inner object is passed to parseTweet along with the extracted card data. This branching logic ensures that legacy payloads are handled separately from newer GraphQL structures while still sharing the same internal Tweet type. The detection gate is located around lines 85‑90 of src/parser.nim.
Mapping Legacy Twitter API Tweet Fields in parseTweet
The parseTweet procedure in src/parser.nim performs the bulk of the field mapping. It accepts the legacy JSON node and constructs a fully populated Tweet object by reading classic attributes such as id_str, full_text, and created_at. It also resolves reply metadata and populates a TweetStats instance with engagement counters.
Key mappings implemented in parseTweet include:
- Tweet ID: extracted from
js{"id_str"}.getId - Text content: read from
js{"full_text"}.getStr - Timestamp: parsed from
created_atorcreated_at_ms - Statistics:
reply_count,retweet_count, andfavorite_countmapped toTweetStats
proc parseTweet(js: JsonNode; ...): Tweet =
result = Tweet(
id: js{"id_str"}.getId,
text: js{"full_text"}.getStr,
time: if js{"created_at"}.notNull:
js{"created_at"}.getTime
else:
js{"created_at_ms"}.getTimeFromMs,
stats: TweetStats(
replies: js{"reply_count"}.getInt,
retweets: js{"retweet_count"}.getInt,
likes: js{"favorite_count"}.getInt
)
)
This procedure spans roughly lines 81‑112 in src/parser.nim and serves as the primary bridge between Twitter's legacy schema and Nitter's internal representation.
Resolving Media from the Legacy Twitter API Tweet Format
After the core Tweet struct is assembled, Nitter processes media attachments by calling parseLegacyMediaEntities on the original JSON node. This procedure walks the extended_entities.media array and creates a Media variant for each entry. The variant distinguishes between photo, video, and animated_gif types based on the type field in the legacy payload.
parseLegacyMediaEntities(js, result)
Inside the procedure, the code iterates over the media array as shown below:
with jsMedia, js{"extended_entities", "media"}:
for m in jsMedia:
case m.getTypeName
of "photo": # create Photo media
of "video": # create Video media with attribution
of "animated_gif": # create Gif media
The resulting Media objects are appended directly to Tweet.media, making them available for timeline and detail rendering. Video entries also carry attribution data when present in the legacy entity. This logic is implemented in lines 107‑133 of src/parser.nim.
End-to-End Legacy Tweet Parsing Example
In practice, the entire flow from raw API response to structured tweet is orchestrated within src/parser.nim. A resolver fetches JSON from an endpoint such as https://api.twitter.com/2/timeline/... and passes it to parseGraphTweet. The function auto-detects the legacy wrapper, strips it away, and delegates to parseTweet and parseLegacyMediaEntities in sequence.
import packedjson, parser
let payload = fetchTwitterAPI("/2/timeline/...") # returns JsonNode
let tweet = parseGraphTweet(payload) # auto-detects legacy
echo tweet.text # full text
echo tweet.stats.likes # like count
echo tweet.media.len # attached media count
You can also extract and inspect media variants after parsing:
let tweet = parseTweet(js{"legacy"})
for m in tweet.media:
case m.kind
of photoMedia: echo "Photo:", m.photo.url
of videoMedia: echo "Video:", m.video.url
of gifMedia: echo "GIF:", m.gif.url
Helper utilities from src/parserutils.nim—including getStr, getId, and select—power the safe JSON navigation throughout these procedures. For GraphQL-specific pathways that occasionally fall back to legacy parsing, see src/experimental/parser/graphql.nim.
Summary
- Nitter detects legacy payloads in
src/parser.nimby checking for a top-level"legacy"key insideparseGraphTweet. - The
parseTweetprocedure maps legacy fields such asid_str,full_text, andfavorite_countto the internalTweetandTweetStatstypes. - Media attachments are resolved by
parseLegacyMediaEntities, which iterates overextended_entities.mediaand producesphoto,video, oranimated_gifvariants. - The entire pipeline is self-contained in
src/parser.nimand relies on helpers fromsrc/parserutils.nimand type definitions fromsrc/types.nim.
Frequently Asked Questions
What file handles the legacy Twitter API tweet parsing in Nitter?
The central parsing logic lives in src/parser.nim. This file contains the parseGraphTweet, parseTweet, and parseLegacyMediaEntities procedures that detect, map, and enrich legacy tweet payloads. Supporting type definitions are located in src/types.nim, and JSON navigation helpers are provided by src/parserutils.nim.
How does Nitter distinguish between legacy and GraphQL tweet payloads?
Inside parseGraphTweet, Nitter checks whether the JSON node contains a "legacy" key or a GraphQL "rest_id" key. If neither is present, it returns an empty Tweet(). When "legacy" is found, the nested object is sent to the legacy parser; otherwise, the GraphQL pathway in src/experimental/parser/graphql.nim takes over.
Which legacy JSON fields does Nitter map to the internal Tweet type?
According to the source code in src/parser.nim, Nitter maps id_str to the tweet ID, full_text to the text content, created_at or created_at_ms to the timestamp, and reply_count, retweet_count, and favorite_count to the TweetStats counters. These mappings occur inside the parseTweet procedure.
How does Nitter handle media attachments from legacy tweets?
Nitter calls parseLegacyMediaEntities after building the core tweet object. This procedure reads the extended_entities.media array from the legacy JSON and generates a Media variant for each attachment. It supports photos, videos, and animated GIFs, and also extracts video attribution when available.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →