Text Cleaner
Fix text that looks fine but behaves strangely. Remove invisible characters and odd spaces, strip HTML and links, straighten curly quotes and tidy blank lines, with a report of exactly what was changed.
- Runs in your browser
- No sign-up
- Free to use
How to use Text Cleaner
- Paste the messy text into the left box.
- Tick the cleaning rules you want; sensible defaults are already selected.
- Read the summary under the text to see what was removed or replaced.
- Copy or download the clean text.
Text Cleaner features
Twelve cleaning rules
Switch each on or off: invisible characters, special spaces, quotes, repeated spaces, trimming, blank lines, HTML, entities, links, emoji, control characters and Unicode normalisation.
Finds what you cannot see
Zero-width spaces, byte-order marks and soft hyphens that break search, sorting and code.
Change report
Counts each kind of fix so you know what was hiding in your text.
HTML to plain text
Strip tags and decode entities such as & and .
ASCII-safe output
Turn typographic quotes, dashes and ellipses into plain characters for code and data files.
Private
Text is cleaned in your browser.
When to use Text Cleaner
- Cleaning text copied from PDFs, Word or web pages before pasting into a CMS.
- Fixing code or data that fails because of invisible characters.
- Preparing text for CSV, JSON or database import.
- Removing links, emoji or markup from user-generated content.
- Making two pieces of text comparable by normalising spaces and Unicode.
Text Cleaner FAQ
What are invisible characters?
Characters with no visible glyph, such as the zero-width space (U+200B), zero-width joiner, byte-order mark and soft hyphen. They are copied along with text from web pages and documents and can break searches, URLs, passwords and code.
What is a non-breaking space?
A space that prevents a line break (U+00A0). It looks like a normal space but is a different character, so text containing it may not match searches or split correctly.
Will removing emoji affect other symbols?
Only emoji and pictographs are removed. Letters, numbers, punctuation and currency symbols stay.
What does normalising Unicode do?
Some characters can be stored in two ways, for example é as one character or as e plus an accent. Normalisation (NFC) converts them to one consistent form so identical-looking text is truly identical.
Why straighten quotes?
Curly quotes are correct in prose but break code, CSV files and command lines, which expect straight ASCII quotes.
The hidden mess in copied text
Text that has travelled through word processors, PDFs, web pages and chat apps accumulates characters that people cannot see but computers treat as significant. A zero-width space in a product code makes it unsearchable. A non-breaking space in a CSV column stops numbers being recognised. A byte-order mark at the start of a file can break a script. These problems are frustrating precisely because the text looks correct.
Word processors also replace straight quotes and hyphens with typographic versions, which is good typography but a common cause of errors when the text is pasted into code, configuration files or spreadsheets.
Unicode adds another subtlety: the same visible character can have more than one representation. An accented letter may be one code point or a base letter followed by a combining accent. The two look identical but do not compare as equal, which matters for searching, sorting and deduplicating. Normalising to a single form removes the difference.
The cleaning rules here are independent, so you can apply exactly what a task needs: gentle tidying for prose, or strict ASCII-safe output for data and code. The report tells you what was found, which is often the first clue to why something was not working.