PDFIntact

Estrai tabelle e testo dal tuo PDF, intatti

Copiare da questo PDF dà caratteri illeggibili

Selezioni il testo in un PDF, lo copi, lo incolli — e ottieni una sequenza illeggibile appartenente a un altro sistema di scrittura. A schermo la pagina è perfetta; si rompe solo nel momento in cui copi.

La causa è che ai font incorporati nel PDF manca la tabella che associa ogni glifo a un codice di carattere — la ToUnicode CMap. Per la visualizzazione basta «disegna questa forma», quindi la pagina appare corretta anche senza. Copiare richiede il codice del carattere, ed è lì che il collegamento si perde. È frequente in articoli scientifici, documenti della pubblica amministrazione e file prodotti da vecchi programmi di impaginazione.

I caratteri non sono rovinati. Sulla pagina sono disegnate forme leggibili. Per questo rileggere le pagine come immagini recupera il testo. Cambiare font o incollare in un altro programma non serve: l'informazione era già persa al momento della copia.

Provalo ora

Verificalo sul tuo PDF

La verifica avviene interamente nel tuo browser. Il file non viene mai caricato e non serve alcun account. Prima di pagare vedi quante pagine, tabelle e figure si possono estrarre.

Esempio

実際のPDFからコピーした文字列

下は官公庁の資料からコピーした実物です。左が化けたもの、右が復元したものです。左をコピーして、お使いのエディタに貼ってみてください。同じ結果になります。

コピーした結果

ᆅ᪉බඹᅋయ࡟࠾ࡅࡿ᭩㠃つไࠊᢲ༳ࠊᑐ㠃つไࡢぢ┤ࡋ࡟ࡘ࠸࡚

PDFIntact で取り出した結果

地方公共団体における書面規制、押印、対面規制の見直しについて

お手元の文字列で判定する

コピーした文字列を貼ると、それがこの種類の文字化けかどうかをその場で判定します。PDFを用意しなくても、コピーした結果だけで確かめられます。

判定はこのブラウザの中だけで行っています。入力した文字列はどこにも送信されません。

Domande frequenti

Why does copying from a PDF produce garbled text?
Because the fonts embedded in the file are missing the table (the ToUnicode CMap) that says which character each glyph shape represents. The page displays correctly, but copying yields different characters. The text is not damaged — only the mapping is missing — so reading the pages as images recovers it.
Can the text be recovered from a garbled PDF?
Yes. PDFIntact reads the pages back as images and recovers the text with character recognition. The check runs entirely in your browser before you pay, so you can confirm it works on your file. Nothing is uploaded.
Will pasting into Word fix it?
No. The character codes are already lost at the point of copying, so the same corruption appears wherever you paste. Changing fonts does not help either.
How much does it cost?
A one-off charge based on page count: 160 JPY for 1–10 pages, 300 for 11–30, 550 for 31–60, 850 for 61–100 and 1,500 for 101–200. No account required. The check itself is free.

Verifica un PDF su tutti i sintomiChe cosa contiene ogni formato