Showing posts with label a32. Capturing Paper Documents. Show all posts
Showing posts with label a32. Capturing Paper Documents. Show all posts

Monday, June 23, 2008

Using the Paper Capture Online Service


Adobe’s Paper Capture feature in Acrobat 6 is designed for individual or small office use. For the needs of larger businesses, Adobe provides their Create Adobe PDF Online service that enables you convert any type of business document to PDF. Company reports, printed archival materials, spreadsheets, calendars, and even entire Web sites are just a few of the types of documents that you can convert in order to take advantage of the universal file-sharing aspects of PDF. The service is subscription based (U.S. $9.99 per month or about U.S. $99 per year), but Adobe offers the service on a trial basis that allows you to create five PDF files free of charge. You can go to Adobe’s Web site and see what all the excitement is about by typing this URL into your favorite browser’s Address text box:
http://createpdf.adobe.com
After you’ve subscribed to the service, you can then upload as many scanned files (of no more than 50 pages in length) as you want and process them online with Paper Capture as follows:
  1. Use your Web browser to go to createpdf.adobe.com, sign in by entering your username and password in the Adobe ID and Password text boxes, and then click the Login button. The Create Adobe PDF page appears.
  2. Click the Choose a File graphic link to open the Create Adobe PDF Online - Select a File dialog box. Note that you can also click the Submit a URL link in order to capture a Web page. A page appears where you specify which file to process.
  3. Click the Browse button to locate the desired file on your hard drive, click Choose, and then click the Continue button on Adobe’s Select a File dialog box to open the Conversion Settings window . Note that you can click the Supported File Types link to view a list of File types supported by the Create Adobe PDF Online service.
  4. Click the Optimization Settings drop-down list and choose either Web (the default), eBook, Screen, Print, or Press as the output conversion setting for your file.
  5. Click the PDF Compatibility drop-down list to select either Acrobat 3.0 (PDF 1.2) (the default), Acrobat 4.0 (PDF 1.3), or Acrobat 5.0 (PDF 1.4) as the output compatibility setting for your file.
  6. Choose a level of security for the converted PDF by clicking the Security Options drop-down list. The default is No Security. You have the option of choosing two other basic levels of security: No Printing (40 bit) or No Printing (128 bit). You can further customize security settings for your converted PDF by clicking the Adobe Acrobat Security link above the Security Options drop-down list.
  7. Select the desired method for having the processed file returned to you in the Delivery Method drop-down list. Your choices are No E-Mail, Download from Conversion History (which lets you archive PDF files at Adobe and download them as necessary from your Conversion History list), Wait for PDF Conversion in Browser, E-Mail Me a Link to My New PDF, or E-Mail Me My New PDF as an Attachment.
  8. Click the Create PDF button at the bottom of the window to upload your file and have it processed according to your wishes.
Create Adobe PDF Online lets you create and save your own conversion settings, just as you would in Acrobat 6. To do so, click the Preferences link under the heading Set Options in the Conversion Settings window and select the new settings using the drop-down lists provided for various conversion settings in the Preferences window. Then click the OK button, enter a descriptive name for the new settings in the dialog box that appears, and click OK. Your new conversion settings will appear in the Optimization Settings drop-down list in the Conversion Settings window.
When the Create Adobe Acrobat Online service receives your uploaded document, it displays a Confirmation screen that gives you an identification number and that indicates how the processed file will be delivered to you. Depending upon your settings, the service then delivers the processed PDF file to you either by displaying it in your Web browser (assuming that you use one that supports the plug-in for displaying PDF files), in an e-mail message as a link or a file attachment, or as a link in your Conversion History list.

Monday, June 16, 2008

Importing Previously Scanned Documents into Acrobat

If you already have a scanned document or an electronic fax saved on your hard drive in a graphics format such as TIFF or BMP (the Tagged Information File Format and Bitmap format are most commonly used for saving scanned images), you can open the file in Acrobat 6 and then process its pages with the Paper Capture plug-in (as described in the previous section). Note that in order for the Paper Capture plug-in to render a searchable PDF document, the source document must be scanned at a resolution setting between 200 and 600 dpi. To open the scanned graphic file in Acrobat, follow these steps:
  1. Choose File➪Create PDF➪From File to display the Open dialog box.
  2. Browse to the folder that contains the graphics file containing the scanned image and click its file icon. If the graphics file is saved in a graphics format other than TIFF, select this file format in the Files of Type drop-down list (the Show drop-down list on the Mac) so that its file icon is displayed in the Open dialog box.
  3. Click the Open button. The scanned graphic is displayed in the Document window in Acrobat.
  4. To save the graphics file as a PDF file, choose File➪Save, and then edit the filename and the folder in which you want to save it (if desired) before clicking the Save button.
  5. To make the text in the new PDF file searchable, choose Document➪ Paper Capture➪Start Capture. The Paper Capture dialog box opens.
  6. To modify the Paper Capture settings before using it to process the pages of your PDF document, click the Edit button to open the Paper Capture Settings dialog box. Otherwise, skip to Step 11.
  7. Select the language of the text in the Primary OCR Language dropdown list.
  8. In the PDF Output Style drop-down list, select one of the following:
    • To be able to both search and edit the text, select the Formatted Text & Graphics option.
    • To make the document text searchable only, select the Searchable Image (Exact) option.
    • To make the text in a document containing many images searchable, select the Searchable Image (Compact) option instead.
  9. To compress the graphics in the PDF document, select the amount of compression in the Downsample Images drop-down list. Your choices are Low (300 dpi), Medium (150 dpi), or High (72 dpi).
  10. Click OK to close the Paper Capture Settings dialog box and return to the Paper Capture dialog box.
  11. Click OK in the Paper Capture dialog box to begin the page processing.
  12. Choose File➪Save a second time to save your changes. After processing the pages of a scanned image that you’ve saved as a PDF document with Paper Capture, if you used the Formatted Text & Graphics output style, you can locate and eliminate all OCR errors in the text by following the steps in the preceding post, “Correcting Paper Capture boo-boos.”

Correcting Paper Capture boo-boos


Although the OCR (Optical Character Recognition) software used by Paper Capture has become better and better over the years, it’s still far from perfect. After processing a scanned PDF document using the Formatted Text & Graphics output style, you need to check your processed document for words that Paper Capture didn’t recognize and therefore wasn’t able to convert from bitmapped graphics into text characters.
To make this check and correct these OCR errors, follow these steps:
  1. Choose Document➪Paper Capture➪Find First OCR Suspect. The program flags the first unrecognized word in the text by putting a gray rectangle around it and opens the Find Element dialog box. Acrobat shows a magnified view of the unrecognized word in the Find Element dialog box,
  2. Choose the TouchUp Text tool by clicking its button on the Advanced Editing toolbar.
  3. In the Find Element dialog box, choose one of the following options:
    • To accept the word displayed and convert it from a graphic into text and then continue to the next capture suspect, click the Accept and Find button.
    • To edit the suspect word directly in the Find Element dialog box, type over incorrect characters in the suspect word and then click the Accept and Find button and go to the next suspect.
    • To ignore an unrecognized word and not convert it to text, just click the Find Next button to move right on to the next suspect.
  4. Repeat Step 3 until you’ve checked and corrected all the unrecognized words in the processed document. Note that if you choose Document➪Paper Capture➪Find All OCR Suspects, the program finds and highlights all suspect elements in the document without opening the Find Element dialog box. This allows you to individually choose which OCR suspect you’d like to edit.
  5. To edit one of the OCR Suspects in a document after choosing Find All OCR Suspects command, make sure the TouchUp Text tool is selected and double-click the desired element to open the Find Element dialog box. The selected OCR Suspect appears in the Find Element dialog box. You can continue by repeating Step 3 or close the Find Element dialog box and repeat Step 5.
  6. Click the Close button in the lower-right corner of the Find Element dialog box to close it, and then choose File➪Save to save your corrections to the PDF document.

How to make scanned documents searchable and editable


When you scan a document directly into a PDF file (as described in the preceding section), Acrobat captures all the text and graphics on each page as though they were all just one big graphic image. This is fine as far as it goes, except that it doesn’t go very far because you can neither edit nor search the PDF document. (As far as Acrobat is concerned, the document doesn’t contain any text to edit or search — it’s just one humongous graphic). That’s where the Paper Capture plug-in in Acrobat 6 for Windows comes into play: You can use it to make a scanned document into a PDF that you can either just search or both search and edit.

To use Paper Capture, all you have to do is choose Document➪Paper Capture to open the Paper Capture dialog box, select the page or pages to be processed (All Pages, Current Page, or From Page x to y), and then click the OK button; the Paper Capture utility does the rest. As it processes the page or pages in the document that you designated, a Paper Capture Plug- In alert dialog box keeps you informed of its progress in preparing and performing the page recognition. When Paper Capture finishes doing the page recognition, this alert dialog box disappears, and you can then save the changes to your PDF document with the File➪Save command When doing the page recognition in a PDF document, the Paper Capture plugin offers you a choice between the following three Output Style options:
  • Searchable Image (Exact): Select this option to make the text in the PDF document searchable but not editable (this is the default setting). This setting is the one to choose if you’re processing a document that needs to be searchable but should never be edited in any way, such as an executed contract.
  • Searchable Image (Compact): Select this option to make the text in the PDF document searchable but not editable and to compress its graphics. Use this setting if you’re processing a document whose text requires searching without editing and that also contains a fair number of graphic images that need compressing. When you select this setting, Paper Capture applies JPEG compression to color images and ZIP compression to black-and-white images.
  • Formatted Text & Graphics: Select this option to make the text in the PDF document both editable and searchable. Pick this setting if you not only want to be able to find text in the document but also possibly make editing changes to it.
To select a different output style setting, click the Edit button in the Paper Capture dialog box to open the Paper Capture Settings dialog box (as shown in Figure 6-5). This dialog box not only enables you to select a new output style in the PDF Output Style drop-down list, but also enables you to designate the primary language used in the text in the Primary OCR Language drop-down list (OCR stands for Optical Character Recognition, which is the kind of software that Paper Capture uses to recognize and convert text captured as a graphic into text that can be searched and edited).

If your PDF document contains graphic images, you can tell Paper Capture how much to compress the images by selecting the maximum resolution in the Downsample Images drop-down list. This menu offers you three options in addition to None (for no compression): Low (300 dpi), Medium (150 dpi), and High (72 dpi). The Low, Medium, and High options refer to the amount of compression applied to the images, and the values 300, 150, and 72 dpi (dots per inch) refer to their resolution and thus their quality. As always, the higher the amount of compression, the smaller the file size and the lower the image quality.

After processing the pages of your PDF document with the Paper Capture plug-in, use the Search feature (Ctrl+F on Windows and Ô+F on the Mac) to search for words or phrases in the text to verify that it can be searched. If you used the Formatted Text & Graphics output style in doing the page recognition, you can select the TouchUp Text Tool by clicking its button on the Advanced Editing toolbar or by typing T, and then click the I-beam pointer in a line of text to select the line with a bounding box to verify that you can edit the text as well. Always remember to choose File➪Save to save the changes made to your document by processing with Paper Capture.