Guide

Invisible characters in AI text

Text copied from an AI tool can carry characters you can't see: zero-width spaces, unusual spaces, soft hyphens, direction controls, and whole messages spelled out in invisible tag characters. They aren't the watermark, but they can mark a text, break searches and code, and smuggle in instructions. glyphwash finds every one, shows you where it was and takes it out.

Remove invisible charactersLast checked

The characters to look for

  • Zero-width characters

    Zero-width space U+200B, zero-width non-joiner U+200C and joiner U+200D, word joiner U+2060, zero-width no-break space U+FEFF, Mongolian vowel separator U+180E, combining grapheme joiner U+034F.

    glyphwash: Removed. Joiners stay where they do a job: inside emoji, and between letters in scripts such as Arabic and Devanagari, where they change how letters connect.

  • Unusual spaces

    No-break space U+00A0, narrow no-break space U+202F, thin and hair spaces U+2009 U+200A, figure and punctuation spaces U+2007 U+2008, en, em and smaller fixed-width spaces U+2002 to U+2006, medium mathematical space U+205F.

    glyphwash: Replaced with an ordinary space.

  • Soft hyphens

    U+00AD, a hyphen that only appears when a word breaks at the end of a line.

    glyphwash: Removed.

  • Text-direction controls

    Direction marks U+200E U+200F U+061C, embeddings and overrides U+202A to U+202E, isolates U+2066 to U+2069.

    glyphwash: Removed. In text with Hebrew or Arabic the marks, embeddings and isolates stay, since they're needed there. Overrides, which can make text display in a different order from how it's stored, never stay.

  • Invisible maths operators

    Function application U+2061, invisible times U+2062, invisible separator U+2063, invisible plus U+2064.

    glyphwash: Removed.

  • Tag characters

    U+E0000 to U+E007F, invisible copies of letters, digits and symbols that can spell out a whole message.

    glyphwash: Removed, and any message they spell out is shown to you. Emoji flags that use them, such as Scotland's, stay.

  • Variation selectors

    U+FE00 to U+FE0F and U+E0100 to U+E01EF, which pick a style for the character before them.

    glyphwash: Removed, except one that styles an emoji or a Chinese or Japanese character. A run of them is how data hides in text, so a run always goes.

  • Look-alike letters

    Letters from other alphabets posing as Latin ones, such as the Greek omicron U+03BF in "repοrt".

    glyphwash: Replaced with the letter they imitate.

Where they come from

  • Copying and pasting. Web pages, PDFs and word processors are full of no-break spaces and soft hyphens, and they come along with the text.
  • Model habits. In April 2025 some answers from OpenAI's o3 and o4-mini contained narrow no-break spaces, which OpenAI said were a quirk of training, not a watermark.
  • Content credentials. The C2PA standard can embed a signed record of where a text came from inside the text itself, written as a zero-width no-break space followed by a run of variation selectors.
  • Fingerprints and hidden messages. Zero-width and tag characters can encode an ID or a message that travels with the text wherever it's pasted, including instructions aimed at AI tools that read it.

Why remove them

  • They can identify a text: the same invisible pattern shows where a copy came from.
  • They break things: searches miss words with a zero-width space inside, code stops compiling, word counts drift.
  • They can carry instructions: a message in tag characters is invisible to you but readable by an AI model.
  • Look-alike letters make words that look right but don't match in searches or spell checks.

Removing them doesn't remove the watermark

The watermarks Claude, ChatGPT and Gemini use aren't characters at all. They're patterns in which words were chosen, so a text with every hidden character removed still carries one. glyphwash takes the characters out first, then rewrites every passage with an open-weight model that adds no watermark of its own. How the Claude watermark works.

Finding them yourself

  • Paste the text into glyphwash: the result lists each hidden character, where it was and what replaced it, and shows any hidden message.
  • Open it in a code editor such as VS Code, which highlights invisible and look-alike characters.
  • In code, this JavaScript pattern finds the invisible ones, though not unusual spaces or look-alike letters:
/[\u00AD\u034F\u061C\u180E\u200B-\u200F\u202A-\u202E\u2060-\u2064\u2066-\u2069\uFEFF\u{E0000}-\u{E007F}]/gu

Questions

Does ChatGPT add invisible characters to its text?

Not as a watermark. In April 2025 some o3 and o4-mini answers contained narrow no-break spaces, which OpenAI said were a quirk of training. OpenAI's watermark, textGrain, is a pattern in word choices instead.

Is a zero-width space a watermark?

Not the kind AI providers use. Zero-width characters can tag or fingerprint a text, but they're easy to strip. The watermarks of Claude, ChatGPT and Gemini are in the choice of words, which survives copying and cleaning.

How do I remove invisible characters from a text?

Paste it into glyphwash, which removes hidden characters, turns unusual spaces into ordinary ones and puts real letters back for look-alikes, then lists everything it changed. A code editor such as VS Code also highlights them, so you can delete them by hand.

Will removing them change how my text looks?

No. They're invisible, unusual spaces become ordinary spaces, and look-alike letters become the letters they were imitating. Characters doing a job, like the joiners inside an emoji, stay.

What is a tag character?

An invisible copy of a letter, digit or symbol from a special block of Unicode. Strung together, tag characters can spell out a whole hidden message. Emoji flags such as Scotland's use them legitimately; in ordinary text they're suspect.

Sources

  1. Unicode: Tags, code chart for U+E0000 to U+E007F, latest version
  2. C2PA: Content Credentials specification 2.4, embedding manifests into unstructured text, 1 April 2026
  3. ITC.ua: The new ChatGPT models leave extra characters in the text, 23 April 2025