What are zero-width spaces and how do I remove them?
Invisible Characters: Find and Remove Them
A zero-width space (U+200B) is a Unicode character that marks a place where a line may break but takes up no width, so it is invisible on screen. It, and a handful of relatives such as the non-breaking space, the byte-order mark and the soft hyphen, get into text through copy and paste from web pages, word processors, chat apps and PDFs. They are harmless in prose and destructive in code, data and identifiers: a variable name, API key or CSV header that contains one looks identical to the correct value and compares unequal. To remove them, strip the zero-width and formatting characters, replace the special spaces with ordinary spaces, and check the result, either with a regular expression in your code or by pasting the text into a cleaner that shows what it removed.
The usual suspects
| Code point | Name | Typical source | What it breaks | Text Cleaner |
|---|---|---|---|---|
| U+200B | Zero width space | CMSs and sites that insert break opportunities in long words or URLs; chat apps | Identifiers, URLs, passwords, search, == comparisons |
Removed |
| U+00A0 | No-break space | in HTML, word processors, Option+Space on a Mac keyboard |
JSON parsing, shell commands copied from documentation, split(" ") |
Replaced with a space |
| U+202F | Narrow no-break space | French and other locales' number grouping (1 234 567,5 from Intl with fr-FR); some ICU versions' time formatting |
Number parsing, comparisons against expected output | Replaced with a space |
| U+2007, U+2009 and others in U+2000–U+200A | Figure space, thin space and similar | Typesetting, PDFs | Same as the no-break space | Replaced with a space |
| U+FEFF | Byte-order mark (zero width no-break space) | Windows tools saving "UTF-8 with BOM"; files concatenated together | JSON, shebang lines, the first CSV header, PHP output | Removed |
| U+00AD | Soft hyphen | Word processors, hyphenation in e-books and web text | Search and matching: co\u00ADoperate does not match cooperate |
Removed |
| U+2060 | Word joiner | Typesetting | Same as the zero width space | Removed |
| U+200C, U+200D | Zero width non-joiner, zero width joiner | Legitimate in Persian, Arabic and Indic scripts, and inside emoji sequences | Only identifiers and ASCII data | Kept inside emoji and next to non-Latin letters, removed elsewhere; see below |
| U+200E, U+200F, U+061C | Left-to-right, right-to-left, Arabic letter marks | Text mixing scripts, copied from bidirectional UIs | Filenames, comparisons | Removed |
| U+202A–U+202E, U+2066–U+2069 | Bidirectional embeddings, overrides and isolates | Mixed-direction text; malicious source code | Code review: text displays in a different order from how it is parsed | Removed |
| U+2028, U+2029 | Line and paragraph separator | Text copied from some editors and PDFs | JavaScript string literals before ES2019; log parsers | Converted to ordinary line breaks |
| U+E0000–U+E007F | Tag characters | Hidden text in prompts and messages; emoji flag sequences | Anything that compares, stores or forwards text | Removed, except inside emoji flag sequences |
| U+000D | Carriage return (the CR in CRLF) | Files saved on Windows | Shell scripts ($'\r': command not found), diffs, trailing-whitespace checks |
Converted with Line endings |
Where each one comes from, and why it matters
Zero width space
Content systems insert U+200B to let long words and URLs wrap; chat and collaboration apps sometimes do the same. When that text is copied into code, a configuration file or a terminal, the character comes along. The symptoms are baffling: an environment variable "does not exist" although it is spelled correctly, a JSON key is undefined, an API rejects a token that looks right. Neither JavaScript's trim() nor Python's strip() removes it, and \s does not match it, because Unicode classifies it as a format character, not as white space.
Non-breaking and narrow spaces
U+00A0 prevents a line break between two words, and word processors insert it after short words or before units. It looks like a space and matches \s in most regex engines, but it is not a space to a JSON parser, a shell, or code that splits on " ". Python shows the difference neatly: "a\u00a0b".split() gives ['a', 'b'], while "a\u00a0b".split(" ") gives one element. Locale-aware formatters emit these characters on purpose: Intl.NumberFormat('fr-FR') groups digits with U+202F, and for a period some ICU versions put U+202F before AM/PM in US English times, which broke code that parsed or compared formatted times.
Byte-order mark
At the start of a UTF-8 file, U+FEFF is a byte-order mark. UTF-8 has no byte order, so it serves only as an encoding signature, and most Unix tools do not expect it. It breaks a shebang line (the kernel reads the BOM bytes before #), makes strict JSON parsers fail at line 1 column 1 (see common JSON errors), and silently renames the first column of a CSV to \ufeffid, so row["id"] is missing. In the middle of a file, it is usually the result of concatenating files that each had one. Behaviour differs by language: JavaScript's trim() and \s treat U+FEFF as white space; Python's strip() and \s do not.
Soft hyphen
U+00AD marks where a word may be hyphenated. It is invisible unless the word breaks at that point, so it survives into databases and search indexes, where co\u00ADoperate (with a soft hyphen) and cooperate are different strings.
Bidirectional controls
Override and isolate characters change the display order of text so that right-to-left scripts render correctly. In source code they can make a line display differently from how the compiler reads it: a comment can appear to end before code that is actually inside it. This technique was published in 2021 as "Trojan Source" (CVE-2021-42574), and most compilers and code hosts now warn about these characters. In code, there is no reason to keep them outside string literals that genuinely need them.
Removing them
Paste the text into the Text Cleaner. This input contains a zero width space inside userid, a no-break space and a narrow no-break space in the price, a byte-order mark before name, and a soft hyphen in cooperate:
userid = 42
price: 1 500 €
name,email
cooperateWith the default options (Remove invisible characters and Normalize spaces on), the output looks the same as the input, which is the point, but it now contains only ordinary characters:
userid = 42
price: 1 500 €
name,email
cooperateThe status line reads Removed 3 invisible characters; replaced 2 non-breaking spaces: the zero width space, BOM and soft hyphen were removed, and the two special spaces were replaced with ordinary ones.
A warning about joiners
U+200C and U+200D are not junk in every context. The zero width non-joiner is required for correct spelling in Persian and changes rendering in several Indic scripts, and the zero width joiner glues emoji sequences together: a family emoji is three person emoji joined by two U+200D characters, and removing them displays three separate faces. The Text Cleaner takes the context into account: with Remove invisible characters on, it keeps a joiner that sits inside an emoji sequence or next to a letter from a non-Latin script such as Persian or Devanagari, and removes joiners everywhere else — between Latin letters, digits, spaces and punctuation, where they only cause mismatches. Persian words and emoji families therefore come through intact, while user\u200Did in an identifier is still cleaned. If you need every joiner removed, including in those scripts, add \u200C\u200D to the character class in the code below.
Removing them in code
Match the specific characters rather than a broad class, so that legitimate text survives:
| Language | Remove zero-width and format characters | Replace special spaces |
|---|---|---|
| JavaScript | s.replace(/[\u200B\u2060\uFEFF\u00AD]/g, '') |
s.replace(/[\u00A0\u2000-\u200A\u202F\u205F\u3000]/g, ' ') |
| Python | re.sub('[\u200b\u2060\ufeff\u00ad]', '', s) |
re.sub('[\u00a0\u2000-\u200a\u202f\u205f\u3000]', ' ', s) |
| Java | s.replaceAll("[\\u200B\\u2060\\uFEFF\\u00AD]", "") |
s.replaceAll("[\\u00A0\\u2000-\\u200A\\u202F\\u205F\\u3000]", " ") |
| PostgreSQL | regexp_replace(s, '[\u200B\u2060\uFEFF\u00AD]', '', 'g') (the regex engine reads \uXXXX itself) |
regexp_replace(s, '[\u00A0\u202F]', ' ', 'g') |
Add the bidi controls (\u202A-\u202E\u2066-\u2069) when cleaning identifiers or code. Unicode also defines a Default_Ignorable_Code_Point property; \p{Default_Ignorable_Code_Point} is supported in JavaScript regexes with the u flag and covers these characters plus many more, including the variation selectors that some emoji need.
To find them rather than remove them, search for anything outside printable ASCII. In most editors, a regex search for [^\x00-\x7F] highlights every non-ASCII character; on the command line, grep -nP '[^\x00-\x7F]' file does the same, and cat -A shows carriage returns as ^M.
Preventing them
- Configure editors to show invisible characters, and to save UTF-8 without a BOM.
- Paste configuration values, keys and commands into a plain-text editor or the Text Cleaner before using them, especially when they come from documents, chat or rendered web pages.
- Validate identifiers at the boundary: an API key, username or environment variable name should match a strict pattern such as
^[A-Za-z0-9_-]+$, which rejects every character in the table above. - Normalise line endings in the repository with
.gitattributes(* text=auto eol=lf) and convert existing files with the Line Ending Converter.