← Back to Blog
Technical Deep Dive2026-10-01·Digital Footprint Health Team

Parsing tweets.js: Emoji, CJK Text and Encoding Traps

tweets.jsX archiveencodingemojidigital footprint

After you unzip an X archive, you will find a file called tweets.js. A lot of people open it, call JSON.parse, and get a syntax error on the first line. The file is not corrupt. It is just not plain JSON: there is an assignment statement wrapped around the array, so the real data sits inside a line of JavaScript. Strip that prefix and what remains is standard data.

The parsing step is the easy part. The trouble starts with the characters. Emoji that render as boxes, CJK text that turns into mojibake, and a word count that is always slightly off all trace back to a small set of causes. This piece walks through the prefix problem, the encoding problem, surrogate pairs, and mixed-script text. If you are still working out what lives in the file, start with the field layout of tweets.js first.

Stripping the prefix in three steps

The first line of tweets.js looks roughly like an assignment: window.YTD.tweets.part0 equals a bracketed array. The exact prefix shifts between archive versions, but the shape does not. You have three ways to handle it, and the right one depends on how many files you are dealing with.

  • Quick look: open it in a text editor, find the first opening bracket, delete everything before it, and save as a json file.
  • Scripted: read the whole string, slice from the first opening bracket to the last closing bracket, then hand the result to a parser.
  • Batch: every data file in the archive follows the same pattern, so write one function that strips the prefix and loops over the folder.

The third option pays off because an archive usually ships several data files: posts, likes, direct messages. Editing them by hand invites misses.

Three symptoms of an encoding problem

When characters break, read the symptom before you reach for a different library. Each symptom points at a different layer.

SymptomLikely causeDirection
CJK text shows as mojibake or question marksRead with the wrong codec, often UTF-8 decoded as Latin-1Specify UTF-8 on read and keep the same codec on write
Emoji render as boxes or blanksThe terminal or font cannot render them; the file is fineCheck code points first, then decide whether the environment is the issue
String length does not match what you seeCode points, UTF-16 units and bytes are being treated as the same thingCount characters by code point throughout

The middle row deserves its own note. Most of the time the file is healthy and the viewer is the problem. Write the content into an HTML file and open it in a browser to settle the question in seconds.

Emoji and surrogate pairs

A single emoji looks like one symbol to your eyes. In UTF-16 it may occupy two units, and in bytes it may take four. Measure the same string with length, with a code-point expansion, and with a byte count, and you get three different numbers. That mismatch is the most common source of confusion in this entire workflow.

The practical damage shows up in truncation. Slice a string by length and you can cut an emoji in half, leaving two invalid halves that render as boxes. To truncate by visible character, expand to code points first and join afterwards.

A cheap way to check

Print the code point of every character in the suspicious string and read the output. Normal characters land in familiar ranges. A stray value between 55296 and 57343 means a surrogate pair was split.

CJK and mixed-script text

Chinese and Japanese do not use spaces between words, so splitting a title or a tag on whitespace does nothing useful. Keyword statistics on CJK content need character-level counting or a dedicated tokenizer. When Latin and CJK text sit in the same post, fix your counting rule before you write the script, or the long-text filter will return the same short sentences over and over.

Punctuation needs the same care. Mixed-script posts often carry both half-width and full-width punctuation, so naive deduplication or matching silently skips part of the set. Normalising punctuation before comparison costs almost nothing and catches a surprising number of misses.

Error to fix reference

Error or symptomFix
Unexpected token w in JSONPrefix not fully stripped; slice from the first opening bracket
Unexpected end of JSON inputTrailing semicolon or extra characters captured; end at the last closing bracket
Parses fine but every field is emptyArchive may be sharded; the payload sits one array level deeper
Emoji count comes out lowCounting by length; switch to code points

What to do once parsing works

Structured data is a starting point, not an outcome. The next layer is almost always classification: by year, by keyword, by how much risk a post carries. All three want a local index so you can query the archive repeatedly without touching the network. That is the main reason to keep archive analysis on your own machine, and the trade-offs are covered in local versus cloud processing.

Once the index exists, deletion has a basis. Which posts stay and which go should follow rules you set, not a single date cut-off. When you are ready to act, the batch cleanup walkthrough sets out a workflow that keeps a reversible record at each step.

Write the definitions down

Two people can count the same archive and get different totals, and the difference is almost always definitional: whether reposts count, whether replies count, whether characters are measured by code point. Put those three choices in a comment at the top of the script, and every later comparison becomes meaningful.

Choosing an output format

Once parsing works, the next decision is the output format. Three options cover most needs. A tabular export is easy to filter by year or keyword in a spreadsheet, which suits manual review. The original nested structure is best for further processing. A local index suits repeated queries and cross-field lookups.

If the archive is going to be kept, store at least one copy in the original nested structure. Field definitions change between versions, and the raw data lets you recalculate later instead of re-exporting.

When fields contain commas

Exporting to a tabular format runs into trouble the moment a post contains commas or line breaks, because both will split the columns. Quote every field and escape the contents rather than joining with plain commas. The error is quiet, and it usually surfaces only after the import when the columns are already misaligned.

Structural drift between archive versions

Archive formats have been revised more than once, so files exported in different years are not identical in shape. Earlier versions shard the posts across files, and field names have been renamed since. When one script has to handle archives from several periods, run a field probe first and confirm the key fields exist before processing anything.

The probe is simple: read the file, print the field names on the first-level objects, and compare against the expected list. Failing loudly beats writing empty values without a warning.

digital-footprint-health.shop offers a free footprint check and a local-first processing workflow. To see which of your public posts deserve attention, start with a quick check from the homepage. For bulk work on old posts, the upload page lists the supported data formats, and plans are laid out on the pricing page.

Frequently Asked Questions

Can tweets.js be parsed with JSON.parse directly?

No. The file opens with an assignment statement that wraps the array. Strip the prefix and keep only the text between the first opening bracket and the last closing bracket, and what remains is valid JSON.

Why does CJK text in my archive turn into mojibake?

The usual cause is a codec mismatch on read, such as decoding a UTF-8 file as Latin-1. Specify UTF-8 on both read and write to recover. If the text was already decoded incorrectly and saved back, re-extract from the original archive.

Why does my post count differ from what the site shows?

It is almost always a definitional gap: whether reposts count, whether replies count, and whether characters are counted by code point or UTF-16 unit. Fix those three rules first, then compare.

Emoji show up as boxes. Is the data broken?

Usually not. A terminal or font that cannot render them will show boxes even when the file is intact. Write the same text into an HTML file and open it in a browser to confirm.

What is the first thing to do after parsing?

Build a local index first, organised by year, keyword and risk level. Deciding what to delete is far easier once you can query the archive, and it beats a blanket date cut-off.

Check your own X/Twitter footprint

Free on-device scan. Your archive never leaves your computer.

Start Free Check

Related Reads

Published on 2026-10-01. Last updated 2026-10-01.