AI Text Cleaner: The Complete Guide (2026)

Asterisks, hashes, curly quotes, em dashes, emojis and characters you cannot see: what each one is, why assistants produce it, and how to remove it without losing the structure worth keeping. Plain text or real HTML, and a workflow you can repeat.

On this page

Cleaning AI text means removing the characters an assistant added for its own display and leaving the words you asked for. Assistants write in Markdown and in the typographic habits of published prose, so a pasted answer carries asterisks, hash marks, curly quotes, em dashes, emojis and sometimes characters you cannot see at all. This guide explains where each of those comes from, what it does to your document when it lands in the wrong place, and how to remove it without losing the structure that was worth keeping.

The short version: keep the meaning, decide what to do with the structure, and delete everything else. The AI Text Cleaner does that in your browser in one paste, and the sections below are the reasoning behind each of its options.

What cleaning AI text actually means

Every chat interface renders what the model writes. The model outputs **important**; the interface shows important. The model outputs ## Summary; you see a heading. That rendering step is invisible while you stay inside the chat window, and it vanishes the moment you copy the text somewhere else. Email composers, CMS fields, spreadsheet cells, form inputs, LinkedIn posts, subject lines and plain-text editors take the characters literally, so the reader sees the asterisks and the hashes.

Cleaning is the work of crossing that boundary deliberately. It is not rewriting, paraphrasing or disguising. A good cleaner changes characters, never words, and it tells you what it changed. It also leaves you a choice: throw the structure away and get plain text, or translate the structure into the format your destination understands, which for Google Docs, Word, Notion and most email clients is HTML.

Why assistants write in Markdown and publisher punctuation

Two forces push models toward the same output. The first is training data. Documentation, README files, wikis, forum posts and chat transcripts are written in Markdown, so a model learns that a clear answer has a heading, a bulleted list and a bolded key term. The second is the product. Chat interfaces instruct the model to answer in Markdown because it is cheap to render and easy to scan when the rendering works, and the model complies.

The punctuation comes from the same place. Books, magazines and quality websites use typographic quotes and em dashes, so the model reproduces them, and it does so at a rate many readers now associate with machine writing. None of this is a defect inside the chat window. It becomes one only at the paste boundary, and it is why a cleaning step belongs in every workflow that moves AI text into another application.

Clean a reply now. Paste ChatGPT, Claude or Gemini output into the AI Text Cleaner: it strips Markdown, straightens quotes, replaces em dashes, removes emojis and lists every hidden character by code point. Plain text or real HTML, in your browser.
Advertisement

The complete inventory of leftovers

Here is what typically survives a copy and paste, where it comes from, and the sensible default for each. The CommonMark specification defines most of the syntax in the first column.

LeftoverSourceDefault action
**bold**, *italic*, ~~struck~~Markdown emphasisRemove the marks, keep the words (or convert to strong, em and s tags in HTML mode)
# Heading, ## SubheadingMarkdown ATX headingsRemove the hashes (or convert to h2 and h3)
- item, * item, 1. itemMarkdown listsNormalize bullets to a hyphen, keep numbering (or convert to ul and ol)
> quoteMarkdown blockquoteRemove the arrow (or convert to blockquote)
`code` and fenced blocksMarkdown codeRemove the backticks and fences, never touch the code inside
[text](url)Markdown linksKeep the text, append the URL in parentheses (or convert to a real link)
| a | b | plus a separator rowMarkdown tablesTab-separated rows for spreadsheets (or a real table)
U+201C U+201D U+2018 U+2019Typographic quotesStraight quotes
U+2014 and spaced U+2013Em dash and spaced en dashComma, hyphen or period, by preference
Pictographs, flags, keycaps, skin tonesEmojiRemove, together with the variation selectors and joiners that build sequences
U+200B, U+200D, U+2060, U+FEFF, U+00AD, U+202F, U+00A0Zero-width and typographic space charactersDelete the zero-width ones, turn the special spaces into ordinary spaces
“Certainly!”, “I hope this helps!”Conversational framingOptional removal of the first and last line

Removing Markdown without losing structure

The naive approach, deleting every asterisk, breaks text that uses asterisks legitimately, such as a footnote marker or a multiplication sign. A careful cleaner works the way a Markdown parser does. It treats a line as a heading only when the hashes are at the start and followed by a space. It treats a pair of asterisks as emphasis only when they wrap text on the same line. It recognizes a list item by its position at the start of a line, and it knows that three or more hyphens alone on a line are a horizontal rule, not a bullet.

Code deserves special care. A fenced block, three backticks on their own line, marks text whose exact characters matter: a shell command, a JSON payload, a regular expression. The right treatment is to remove the fences and leave every character between them alone, including asterisks and hashes that would be formatting anywhere else. Inline code in single backticks gets the same protection for its contents.

Links are the one place where stripping loses information. [our pricing page](https://example.com/pricing) collapses to “our pricing page” if you keep only the text, and the reader loses the address. The cleaner keeps the text and adds the URL in parentheses when the two differ, so nothing disappears; in HTML mode it becomes a real link instead.

Tables: from pipes to tabs, or to real HTML

A Markdown table is a header row, a separator row made of hyphens and optional colons, and body rows, all delimited by pipes. Pasted into a spreadsheet it becomes one cell of pipe characters. The plain-text fix is to drop the separator row and turn each pipe into a tab, because spreadsheets split on tabs when you paste. Google Sheets, Excel and Numbers all do this, so a cleaned table lands in a grid of cells with no further work.

When the destination is a document rather than a spreadsheet, HTML mode is better: the table becomes a real table element with a header row, and Docs or Word will keep it as a table on paste. Either way, check the result once. Markdown allows a table cell to contain formatting, and a cell that contained a pipe inside code will have confused the assistant as much as it confuses the cleaner.

Em dashes, en dashes and hyphens

Three characters look similar and behave differently. The hyphen (U+002D) joins words and breaks lines. The en dash (U+2013) marks ranges, as in 2019–2024, and in some styles a spaced en dash stands in for an em dash. The em dash (U+2014) sets off a phrase or joins two independent clauses. Purdue OWL summarizes the usage in its dash guide, and the two main American style guides differ on spacing: the Chicago Manual of Style sets em dashes closed, with no space on either side, while the Associated Press Stylebook puts a space on each side.

AI text uses the em dash heavily, and many editors now read it as a signal of machine drafting. Whether or not that concerns you, there are practical reasons to convert. Some CMS fields and legacy databases mangle non-ASCII punctuation. Subject lines and SMS have character budgets where a comma is cheaper. Plain-text email clients may render the dash as a box. The cleaner offers three replacements. A comma is the most common function an em dash performs and reads naturally in most sentences. A spaced hyphen keeps the visual pause for informal writing. A period, with the next word capitalized, is right when the dash is really joining two sentences that stand on their own.

One rule protects the exceptions: an unspaced en dash between digits is a range and must stay. Converting 1990–2000 to “1990, 2000” would change the meaning, so the cleaner leaves it alone.

Smart quotes and straight quotes

Typographic quotes (U+201C and U+201D for double, U+2018 and U+2019 for single) are correct in print and on most web pages. They are wrong in every context that treats quotes as syntax. JSON requires straight double quotes; a curly quote makes the file unparsable. Shell commands, CSV fields, regular expressions, configuration files and search queries all expect the plain ASCII characters U+0022 and U+0027. Straightening quotes is therefore a near-universal safety measure when AI text goes anywhere near code or data.

The apostrophe deserves a note. English possessives and contractions use the same U+2019 character as a closing single quote, and the straight replacement is the ASCII apostrophe, which is what keyboards produce and what software expects. Guillemets (U+00AB and U+00BB), common in Spanish, French and German typography, become straight double quotes as well. If your destination is a formatted document and you want curly quotes back, most word processors re-curl straight quotes as you type, and a find-and-replace pass restores them in bulk.

Emojis, variation selectors and joiner sequences

Emoji are more than single characters. A flag is two regional indicator symbols. A family or a profession is several emoji glued together with the zero-width joiner (U+200D). A skin tone is a base emoji plus a modifier. And many emoji end with variation selector-16 (U+FE0F), an invisible character that asks for the colorful presentation instead of the text one. Unicode Technical Standard #51 defines all of these mechanisms.

Deleting only the visible pictograph leaves the invisible parts behind, which is how a “clean” text ends up with stray joiners and selectors in it. A correct emoji filter removes the whole sequence: the pictographs and regional indicators, the joiners between them, the skin-tone modifiers and the variation selectors. It should also tidy the space it leaves, so “done ✅.” becomes “done.” rather than “done .”

Hidden characters: the complete list, and the April 2025 episode

Unicode includes characters with no visible glyph. The Unicode Standard describes the format characters and special areas in its chapter on special areas and format characters, and the Unicode Character Database assigns them the general category Cf (format). Some are legitimate typography: the no-break space (U+00A0) keeps a number and its unit on one line, and the narrow no-break space (U+202F) is the correct thin space before certain punctuation in French and inside some numeric conventions. Others are controls: the zero-width joiner and non-joiner shape scripts such as Arabic and Devanagari and build emoji sequences; the byte order mark (U+FEFF) signals encoding at the top of a file; the bidirectional controls (U+202A to U+202E, U+2066 to U+2069) set text direction; the soft hyphen (U+00AD) marks an optional break point; the word joiner (U+2060) forbids one.

In ordinary prose pasted into ordinary software these characters are noise, and the noise has consequences. Search fails to match a word that contains a zero-width space, because to the computer it is a different word. A spreadsheet refuses to treat “1 000” as a number. A password or a license key copied from a chat stops working. A translation memory or a plagiarism checker misses a match. Because the characters are invisible, the only reliable way to catch them is a tool that lists them by code point.

The episode that made this mainstream came in April 2025. On April 20 the education company Rumi published a report that longer responses from OpenAI’s o3 and o4-mini models contained narrow no-break spaces (U+202F) in place of ordinary spaces, and asked whether the pattern was a watermark. On April 22 Rumi added OpenAI’s response: the characters were not a watermark but, in OpenAI’s words, “a quirk of large-scale reinforcement learning.” On April 23 Rumi reported that the characters had disappeared from new responses. Whatever the cause, the lesson stands: AI text can carry characters you did not ask for and cannot see, and a detector that names them is part of any honest cleaning step.

Filler openers and closers

Assistants frame answers conversationally. “Certainly! Here is a summary of…” opens, and “I hope this helps! Let me know if you would like…” closes. Inside a chat that framing is polite. Inside a report, a product description or a knowledge-base article it is a tell, and it is dead weight. Removing it is a judgment call, which is why the cleaner keeps the option off by default and limits it to a stock first line and a stock last line. On a full reply that is exactly right; on an excerpt that happens to begin with “Here is” it could remove a legitimate sentence, so read the output before you trust the option.

Plain text or rich text: choose by destination

The destination decides the mode. If it cannot render formatting, or you do not want formatting there, use plain text: email subject lines, SMS, form fields, spreadsheet cells, code comments, social posts, chat messages, CSV files. If it renders formatting and you want to keep the assistant’s structure, use HTML and copy as rich text. Modern browsers expose a clipboard API that can carry an HTML version and a plain-text version of the same content at once (MDN documents the ClipboardItem interface), so when you paste into Docs or Word you get headings, lists and a table, and when you paste into a plain-text field you get clean text. The cleaner’s HTML mode uses exactly that mechanism.

A middle path exists for LinkedIn and other platforms that support no formatting at all. The bold you see in some posts is not formatting but Unicode mathematical letters that look bold, produced by tools such as the Bold Text Generator. Use that sparingly: those letters are read aloud strangely by screen readers and are invisible to search.

A repeatable workflow

  1. Copy the whole reply, including any opener and closer, so the filler option can recognize them.
  2. Paste into the AI Text Cleaner and read the hidden-character panel first. If it lists anything, note what and how many; a spike in a particular code point is worth knowing about.
  3. Pick the mode. Plain text for fields and messages; HTML for documents.
  4. Set the dash policy to match your house style, then leave the other defaults on.
  5. Copy the output and paste it where it is going. Read it once in its new home, because a cleaner changes characters and not sense, and a sentence written around a dramatic dash may need a small edit once the dash is a comma.
  6. Count what matters. If the text is going back into a prompt, the Token Counter shows how much the cleanup saved; if it is going into a form, the Character Counter confirms it fits.

Cleaning formatting is not the same as hiding authorship. If a publisher, a school or an employer asks you to disclose AI assistance, removing asterisks does not change that obligation, and no formatting cleaner turns machine-drafted prose into human prose. What cleaning does is make the text usable, searchable and safe to store, which is reason enough to make it a habit.

The cleaner covers the common cases with fixed rules. For whitespace-only problems, the Remove Extra Spaces tool collapses double spaces and stray line breaks. For a substitution the presets do not cover, the Find & Replace tool accepts regular expressions. If the AI output was JSON rather than prose, the JSON Formatter validates and pretty-prints it and will flag the curly quote that breaks it. And if you are comparing the assistant’s draft with your edit, the Text Diff tool shows exactly which characters changed.

Sources and further reading

Frequently asked questions

How do I remove ChatGPT formatting before pasting into Google Docs?

Paste the reply into the AI Text Cleaner, choose HTML (keep formatting) and click Copy as rich text. Docs receives real headings, bold, lists and tables instead of asterisks and hashes. If you want no formatting at all, stay in plain text and use Copy output.

Why does ChatGPT use so many em dashes?

Because the prose it learned from, books, magazines and quality websites, uses them, and the model reproduces the habit. Converting them to commas, hyphens or periods is a style choice; the cleaner offers all three and keeps numeric en-dash ranges such as 2019–2024 intact.

Are the hidden characters in ChatGPT text a watermark?

In April 2025 Rumi reported narrow no-break spaces (U+202F) in o3 and o4-mini responses. OpenAI said they were not a watermark but a quirk of large-scale reinforcement learning, and the pattern disappeared within days. Hidden characters still appear in AI and copied text for other reasons, so a detector remains useful.

Does removing formatting make AI text undetectable?

No. A formatting cleaner changes characters, not writing style, and it does nothing to alter an obligation to disclose AI assistance. Its purpose is to make text usable, searchable and safe to store.

Can I keep bold and headings but lose the symbols?

Yes. HTML mode converts Markdown into real HTML elements, and Copy as rich text pastes them as formatted text into Docs, Word, Notion and most email composers, while a plain-text fallback goes to fields that cannot render formatting.

Is it safe to paste confidential text into an online cleaner?

The TextKit cleaner runs entirely in your browser: nothing is uploaded, logged or stored, and closing the tab discards the text. Check that any other tool makes and keeps the same promise before pasting anything sensitive.

Advertisement

Keep reading

Written by . We build the tools we write about. Try the AI Text Cleaner used in this post.