Text Tools

Text Cleaner

Fix text that looks fine but behaves strangely. Remove invisible characters and odd spaces, strip HTML and links, straighten curly quotes and tidy blank lines, with a report of exactly what was changed.

  • Runs in your browser
  • No sign-up
  • Free to use
Cleaning rules

How to use Text Cleaner

  1. Paste the messy text into the left box.
  2. Tick the cleaning rules you want; sensible defaults are already selected.
  3. Read the summary under the text to see what was removed or replaced.
  4. Copy or download the clean text.

Text Cleaner features

Twelve cleaning rules

Switch each on or off: invisible characters, special spaces, quotes, repeated spaces, trimming, blank lines, HTML, entities, links, emoji, control characters and Unicode normalisation.

Finds what you cannot see

Zero-width spaces, byte-order marks and soft hyphens that break search, sorting and code.

Change report

Counts each kind of fix so you know what was hiding in your text.

HTML to plain text

Strip tags and decode entities such as & and  .

ASCII-safe output

Turn typographic quotes, dashes and ellipses into plain characters for code and data files.

Private

Text is cleaned in your browser.

When to use Text Cleaner

  • Cleaning text copied from PDFs, Word or web pages before pasting into a CMS.
  • Fixing code or data that fails because of invisible characters.
  • Preparing text for CSV, JSON or database import.
  • Removing links, emoji or markup from user-generated content.
  • Making two pieces of text comparable by normalising spaces and Unicode.

Text Cleaner FAQ

What are invisible characters?

Characters with no visible glyph, such as the zero-width space (U+200B), zero-width joiner, byte-order mark and soft hyphen. They are copied along with text from web pages and documents and can break searches, URLs, passwords and code.

What is a non-breaking space?

A space that prevents a line break (U+00A0). It looks like a normal space but is a different character, so text containing it may not match searches or split correctly.

Will removing emoji affect other symbols?

Only emoji and pictographs are removed. Letters, numbers, punctuation and currency symbols stay.

What does normalising Unicode do?

Some characters can be stored in two ways, for example é as one character or as e plus an accent. Normalisation (NFC) converts them to one consistent form so identical-looking text is truly identical.

Why straighten quotes?

Curly quotes are correct in prose but break code, CSV files and command lines, which expect straight ASCII quotes.

The hidden mess in copied text

Text that has travelled through word processors, PDFs, web pages and chat apps accumulates characters that people cannot see but computers treat as significant. A zero-width space in a product code makes it unsearchable. A non-breaking space in a CSV column stops numbers being recognised. A byte-order mark at the start of a file can break a script. These problems are frustrating precisely because the text looks correct.

Word processors also replace straight quotes and hyphens with typographic versions, which is good typography but a common cause of errors when the text is pasted into code, configuration files or spreadsheets.

Unicode adds another subtlety: the same visible character can have more than one representation. An accented letter may be one code point or a base letter followed by a combining accent. The two look identical but do not compare as equal, which matters for searching, sorting and deduplicating. Normalising to a single form removes the difference.

The cleaning rules here are independent, so you can apply exactly what a task needs: gentle tidying for prose, or strict ASCII-safe output for data and code. The report tells you what was found, which is often the first clue to why something was not working.

Other useful tools