X's character limit is a budget, not a length
Many tools that count characters for X count them wrong, including some large ones. The reason is that text.length is an obvious thing to reach for and is not what X measures.
The rule
X does not count characters. It counts weighted characters, and the weights are not all one.
Characters in a small set of ranges weigh 1: Latin letters, digits and common punctuation, and also Greek, Cyrillic, Hebrew, Arabic, Devanagari and Thai. Everything else weighs 2. That includes emoji, and it includes Chinese, Japanese and Korean.
The consequence is that the same sentence has a different cost depending on the language it is written in, and a post that fits in English may not fit in Japanese. If you are publishing in more than one language, this is not a detail.
URLs cost a flat 23
Every URL counts as 23 characters, no matter its actual length. A nine-character link and a ninety-character link cost exactly the same, because X shortens both through its own wrapper before measuring.
Two things follow. Shortening a URL before posting saves you nothing at all — the character cost is identical and you have made the link less trustworthy-looking for no gain. And a post that looks 40 characters over the limit in your editor may be comfortably under it, if the overrun is a long URL.
Why tools get it wrong
Because the naive implementation is one property access, and it is right often enough to ship. An all-ASCII post with no links gives the same answer either way, and most posts during development are all-ASCII posts with no links.
It breaks on exactly the posts that matter: the one with an emoji in the hook, the one with a link to the thing you are announcing, the one written in the language most of your audience reads.
What correct looks like
Count code points, not UTF-16 units — an emoji is frequently two units and one character. Assign weight 1 to the light ranges and 2 to everything else. Replace each URL with a flat 23 before counting the remainder. Compare the total against 280.
And then — this is the part that is easy to get wrong even after the arithmetic is right — make sure the number you show the user is the number you actually send. A counter that measures the text in the editor while the publishing code appends hashtags on top is not a rounding error. It is two different notions of "the post", and the user only finds out which one was real after it is public.
We shipped that bug, found it while publishing, and fixed it by making one function the single source for both the counter and the publisher. The test that guards it runs the browser's copy of the rule and the server's copy over the same inputs and fails the build if they ever disagree.
Correction, 27 August 2026: code points are not enough
The paragraph above says to count code points rather than UTF-16 units. That is a real improvement over .length, and it is still not the rule. We know because we followed our own advice and it was wrong.
A four-person family emoji is four people joined by three invisible characters — seven code points. Counting code points and weighting them gives 11. X charges 2. A flag emoji is two letters in disguise; code points give 4, X charges 2. X's configuration says it plainly, next to the weights we had already implemented correctly: the weighted length "considers all emoji as a single code point … including longer grapheme clusters combined by zero-width joiners."
So the rule is grapheme clusters, with any complete emoji charged as one unit at weight 2 — not code points.
The instructive part is how it survived. We had a test comparing the browser's copy of the counter against the server's, over fixtures that included a zero-width-joiner emoji, and it passed every run — because both copies were wrong in the same way. Two implementations agreeing with each other is evidence of consistency, not of correctness. The bug was found only by testing the shipped counter against X's published configuration instead of against a second copy of our own assumption.
Both kinds of test are worth having, and neither substitutes for the other. One stops your copies diverging. The other stops them being confidently wrong together.
The error ran in the direction of over-counting, so nothing unpostable was ever sent — we refused posts X would have accepted. That is still a failure, and it landed on the one feature we describe as showing the real number.
Correction, 4 October 2026: we were deciding for ourselves what a link is
Three things in this post were wrong, and two of them were in our product as well.
The scripts. This post said Arabic, Cyrillic and Devanagari weigh 2. They weigh 1: the cheap ranges run up to U+10FF, which takes in far more than Latin. The paragraph near the top is corrected. Our counter always had this right. The sentence did not.
The links. "Replace each URL with a flat 23" skips the hard part, which is deciding what a URL is. We decided for ourselves, and X decides differently:
- The full stop that ends a sentence is not part of the link before it.
Read more: https://example.com/post.costs 35, not 34. Our counter took the stop into the link, so a post that was one over the limit read as fitting. - A bare domain is a link on any of 1,581 endings, not the 49 we had listed.
whitehouse.govcosts 23, and so doesmain.rs.README.mddoes not: x.com never treats a bare.mdas a link, though the library X publishes does. We found that one on 9 October, by pasting it into X's own composer. - An email address is not a link. We charged
jane@example.com23 for its domain.
We found it by running twitter-text, the library X publishes, over the same strings. Nothing of ours could have found it. By then we had four copies of the counter and 49 published test vectors, all written by the same person, and all of it agreed. The counter is now a port of X's own patterns, and it runs X's own test cases as well as ours.
"Almost every tool." That is how this post used to open, and it was a guess. We have since measured six counters, ours among them, on four test strings. Three were wrong. Ours passed, and was wrong about links. So the opening now says "many", and the list of counters we know to have been wrong has our own on it more than once.
The first correction said that a test against a second copy of your own assumption cannot catch a mistake both copies share. This is the same lesson one level up: a test suite you wrote yourself is also a copy of your own assumption.
If you want the corrected rule without implementing it: the X character counter runs it, free and signed out, and so does the thread splitter. 74 test vectors say what any counter should answer, ours included. Every platform's limits are on one page, marked by whether the platform documents the figure or people merely observe it.
280 is a budget. Spend it deliberately.