1スキャンした PDF documents をアップロードします。クリーンな 300 DPI グレースケールスキャンは認識機能にとって最も有用です。
2言語を指定すると、 言語の名前を付けることで、 推測をさせることよりも 多くの利点が得られます。
3認識パスを実行します。単語は元のページ画像の上に見えないレイヤーとして書き戻されます。
4検索可能な PDF ファイルをダウンロードします。同じように見えますが、検索とテキストの選択が可能です。
OCR PDF よくある質問
ソースフォーマットは認識に影響するのか?
+
It does. PDF はページを固定します: フォント、ベクトル、ラスター画像、テキスト座標は凍結され、読者全員が同じレイアウトを見るようになります. How the page is stored decides what resolution and colour information the recogniser has to work with.
PDF に特有な何かは?
+
Yes — PDFはページの画像ではなく オブジェクトグラフです テキストは選択可能で ベクトルは鋭く残っています 内部のラスターコンテンツに何が起こっても. It affects what the recogniser can see.
Concretely, 各ページはレンダリングされ テッセラクトのテキスト認識を通して 100以上の言語で実行され 認識された単語は 原画像の上に 見えないテキストレイヤーとして書き戻されます ページは変わらないように見えますが 検索可能で選択可能です. The page still looks exactly as it did — the recognised text sits invisibly behind the image so search and selection work without changing the appearance.
WORD.toは文書の編集可能な終わりを中心に構築されており、誰かが送信するために平坦化する前に、まだ書かれ、まだスタイル化され、まだ論争されているDOCXである。 Office files are containers full of other people's media — images, embedded audio, fonts — so the work people need on them is usually the work they would need on those contents anyway. OCR PDF shares the upload, the caps and the account with the conversions for that reason.
OCR PDFが終わったら、結果をどうするか。
+
コンバータは送信用のPDF、埋め込み用の画像、プログラム的に読み込むための普通のテキストに変換する。 Doing that afterwards keeps the editable original around, which is the part you cannot get back once it has been flattened.