Markdown · Em dashes · Hidden characters

AI Text Cleaner

Paste text from ChatGPT, Claude or Gemini. Get clean plain text or real HTML, with every hidden character listed.

0Characters out
0Characters in
0Hidden chars
0Emojis removed
Advertisement

What the AI Text Cleaner does

Paste text from ChatGPT, Claude, Gemini, Copilot or any other assistant and the cleaner gives you back the same words without the formatting baggage. Assistants write in Markdown, a plain-text markup where **two asterisks** mean bold, a leading # means a heading and a leading - means a bullet. Chat interfaces render those marks into styling, but the moment you paste the text into an email, a CMS field, a spreadsheet cell, a LinkedIn post or a plain-text form, the marks come along as literal characters. The cleaner strips them.

It also handles the less visible leftovers. Typographic quotes are turned back into straight quotes, em dashes are replaced with the punctuation you choose, emojis are removed, and hidden Unicode characters, the zero-width spaces, word joiners, soft hyphens and narrow no-break spaces that sometimes ride along in AI output, are found, listed by code point and deleted. Everything runs in your browser. The text is never uploaded.

Two output modes: plain text or HTML

Plain text is the default. Markdown marks are removed, tables become tab-separated rows you can paste straight into a spreadsheet, and bullets are normalized to a simple hyphen. Use it for email, forms, social posts, subject lines, chat messages and anywhere formatting is either impossible or unwanted.

HTML (keep formatting) converts the Markdown into real HTML: headings become <h2> and <h3>, bold becomes <strong>, lists become <ul> and <ol>, pipe tables become <table>. The output box shows the source, and the Copy as rich text button places both the HTML and a plain-text fallback on the clipboard, so Google Docs, Word, Notion and most email composers paste it as formatted text. This is the mode to use when you want to keep the structure the assistant produced and lose only the raw symbols.

Every option, explained

OptionWhat it changesWhat it leaves alone
Remove Markdown marksBold, italic and strikethrough markers, heading hashes, blockquote arrows, inline backticks and code fences, link syntax (the link text stays, followed by the URL in parentheses when they differ), image syntax, horizontal rules, task-list checkboxes, table pipesThe words themselves, line breaks, numbered lists, the content of code blocks
Em dashesSpaced em dashes and spaced en dashes become a comma, a hyphen or a period (with the next word capitalized), or stay as they areUnspaced en dashes between digits, such as 2019–2024, which are ranges
Straighten smart quotesCurly double and single quotes, low-9 quotes, prime marks and guillemets become " and 'Apostrophes inside words are kept as straight apostrophes
Remove emojisPictographs, flags, keycap sequences, skin-tone modifiers and joined emoji sequences, plus the variation selectors that travel with themOrdinary punctuation and symbols such as % and &
Remove hidden charactersZero-width space, non-joiner and joiner, word joiner, byte order mark, soft hyphen, bidirectional controls, invisible math operators and the whole family of typographic spaces (no-break, narrow no-break, thin, hair, em, en, figure)Regular spaces, tabs and line breaks
Remove AI filler linesA first line that is a stock opener (“Certainly!”, “Great question”, “Here is…”) and a last line that is a stock closer (“I hope this helps”, “Let me know if…”). Off by default, because it removes whole linesEverything between them
Normalize spaces and blank linesTrailing spaces, runs of spaces, and three or more blank lines in a row (collapsed to one)Single blank lines between paragraphs

Why AI text carries so many stray symbols

Large language models are trained on enormous amounts of Markdown: documentation, README files, forum posts, chat transcripts. Markdown is also what chat products ask the model to produce, because it is cheap to render and easy to read when the rendering works. So the model learns that a helpful answer has a heading, a bulleted list, a bolded key term and a friendly opener and closer. None of that is wrong inside the chat window. It becomes a problem only at the paste boundary, where the receiving application takes the text literally.

The same training data explains the punctuation habits. Published prose uses typographic quotes and em dashes, so the model reproduces them at a rate many readers now associate with machine writing. Whether or not you want to disguise the origin of a draft, straight quotes and simpler punctuation are what most style guides, code editors, CMS fields and legal templates expect.

Hidden characters: what they are and why they matter

Unicode defines a number of characters that have no visible glyph. Some are legitimate typography: the no-break space keeps “10 kg” from splitting across a line, and the narrow no-break space is the correct thin space in French punctuation and in some numeric conventions. Others are formatting controls: the zero-width joiner glues emoji into a single sequence, the byte order mark tells software how a file is encoded, and the bidirectional controls set text direction. In ordinary prose pasted into ordinary software, all of them are noise, and some of them cause real trouble: search fails to match words that look identical, spreadsheets refuse to treat a “number” as a number, and passwords or codes copied from a chat stop working.

In April 2025 the education company Rumi reported that responses from OpenAI’s o3 and o4-mini models contained narrow no-break spaces (U+202F) where ordinary spaces belonged, which set off a debate about whether the characters were a deliberate watermark. OpenAI told Rumi the characters were a side effect of training rather than a watermark, and the pattern disappeared shortly afterwards. The lesson holds regardless of intent: you cannot see these characters, so a tool has to find them for you. The detector panel above lists every hidden character it finds, with its Unicode code point and a count, before the cleaner removes it.

How to use it well

Start with the defaults. They remove Markdown, straighten quotes, replace em dashes with commas, drop emojis, delete hidden characters and tidy spacing, which is the right combination for email, forms and most publishing tools. Switch Em dashes to “to hyphen” if your house style uses spaced hyphens, or to “to period” when the dashes are joining sentences that read better apart. Turn on Remove AI filler lines only when the text is a full assistant reply with a greeting and a sign-off; on excerpts it can trim a legitimate first line.

When the destination renders formatting, choose HTML (keep formatting) and use Copy as rich text. When it does not, stay in plain text. Then read the output once. A cleaner changes characters, not meaning, and the sentence the assistant wrote with a flourish may need a human edit once the flourish is gone. Pair it with the Remove Extra Spaces tool for whitespace-only cleanup, the Find & Replace tool for a custom substitution the presets do not cover, and the Token Counter if the cleaned text is going back into a prompt.

Privacy and limits

The cleaner is a few hundred lines of JavaScript that run on your device. There is no server, no account and no log; closing the tab discards the text. It is deterministic and rule-based: it does not rewrite, paraphrase or “humanize” anything, and it cannot judge whether a sentence sounds machine-written. It also does not touch the content of fenced code blocks beyond removing the fences, because code often depends on its exact characters. If your text is code, use the JSON Formatter or leave the fences in place.

Frequently asked questions

Does the cleaner remove asterisks from ChatGPT text?

Yes. Double asterisks (bold), single asterisks and underscores (italic), triple asterisks, tildes (strikethrough), backticks and heading hashes are all removed in plain-text mode. Bullet asterisks are normalized to a simple hyphen so the list structure survives.

What is the difference between plain text and HTML mode?

Plain text strips every Markdown mark and turns tables into tab-separated rows. HTML mode converts the Markdown into real headings, lists, links and tables, and the Copy as rich text button puts formatted text on the clipboard for Google Docs, Word, Notion or email.

Which hidden characters does it detect?

Zero-width space, zero-width non-joiner and joiner, word joiner, byte order mark, soft hyphen, combining grapheme joiner, the bidirectional control characters, the invisible math operators, variation selectors, and every typographic space from the no-break space to the ideographic space. Each one is listed with its Unicode code point and a count before it is removed.

Why do em dashes get replaced with commas?

Because a comma is the most common punctuation an em dash stands in for, and it is the default most editors expect. You can switch to a spaced hyphen, to a period with the next word capitalized, or keep the dashes. Unspaced en dashes between numbers, such as 2019–2024, are kept as ranges.

Will it change my words or rewrite the text?

No. The cleaner only changes formatting characters and punctuation marks according to the options you pick. It never paraphrases, reorders or removes sentences, except the optional filler-line setting, which removes a stock opening or closing line and is off by default.

Is my text uploaded anywhere?

No. Everything runs in your browser with JavaScript. Nothing you paste is sent to a server, logged or stored, and closing the tab discards it.

Related

Advertisement

Learn more about cleaning AI text