OCR.space connects to a Nagent workspace with an API key. Once it is connected, agents can call 3 OCR.space actions: "Get Conversion Statistics", "Extract Text from Image/PDF (OCR)" and "Extract Text from Image URL (GET)". Nothing is enabled on connect: each action is allowed one at a time, and an action with side effects runs or waits for a person according to the agent's level.
Every operation an agent can call against OCR.space, with input parameters and output schema.
OCRSPACE_GET_CONVERSIONSRetrieve OCR API conversion statistics and usage data (PRO accounts only). Returns the number of conversions for Engine1, Engine2, and total conversions. Data is updated once daily and shows conversions from start of month to end of yesterday. Free API keys will return 0 conversions.
Input parameters
Start date option for conversion statistics. Case-sensitive.
Output
Data from the action execution
Error if any occurred during the execution of the action
Whether or not the action execution was successful or not
OCRSPACE_OCR_PARSE_IMAGE_POSTExtract text from images and PDF documents using OCR (Optical Character Recognition). Supports 27 languages, table recognition, orientation detection, and word-level coordinate extraction. Provide exactly one of `file`, `url`, or `base64Image`; providing multiple or none triggers E301/OCRExitCode 99. Input can be provided as file upload, public URL, or base64-encoded data URI. Response is nested JSON; extract text from `ParsedResults\[*\].ParsedText`. Returns extracted text with optional overlay coordinates and searchable PDF generation. For poor-quality scans, enable both `detectOrientation` and `scale` and ensure `language` matches the document.
Input parameters
Publicly accessible URL of the image or PDF file. IMPORTANT: Must also set 'filetype' parameter (e.g., 'PNG', 'JPG', 'PDF') when using this option.
Binary content of image or PDF file for upload (multipart/form-data). Supports JPG, PNG, GIF, PDF, BMP, TIF formats.
If True, automatically upscales low-resolution images before OCR processing to improve text recognition accuracy.
If True, optimizes OCR for table/structured data recognition. Recommended for receipts, invoices, and tabular documents.
File format specification. REQUIRED when using 'url' or 'base64Image' parameters. Valid values: 'PDF', 'GIF', 'PNG', 'JPG', 'TIF', 'BMP'. Not needed for 'file' parameter.
OCR language code: ara=Arabic, bul=Bulgarian, chs=Chinese Simplified, cht=Chinese Traditional, hrv=Croatian, cze=Czech, dan=Danish, dut=Dutch, eng=English, fin=Finnish, fre=French, ger=German, gre=Greek, hun=Hungarian, kor=Korean, ita=Italian, jpn=Japanese, pol=Polish, por=Portuguese, rus=Russian, slv=Slovenian, spa=Spanish, swe=Swedish, tha=Thai, tur=Turkish, ukr=Ukrainian, vnm=Vietnamese
OCR processing engine selection. 1=Standard engine (default, reliable), 2=Experimental engine (may have better accuracy for some documents)
Base64-encoded image as a data URI string. Format: 'data:image/\[format\];base64,\[encoded-data\]'. IMPORTANT: Must also set 'filetype' parameter when using this option.
If True, automatically detects and corrects text orientation (rotation). Returns detected orientation angle in TextOrientation field.
If True, returns word-level bounding box coordinates (Left, Top, Height, Width) for each detected word. Useful for document layout analysis.
If True, generates a searchable PDF with an invisible text layer overlay. Returns PDF URL in SearchablePDFURL field.
If True (and isCreateSearchablePdf=True), hides the text layer in the generated searchable PDF. The text is still searchable but not visible.
Output
Data from the action execution
Error if any occurred during the execution of the action
Whether or not the action execution was successful or not
OCRSPACE_PARSE_IMAGE_URLExtract text from images via URL using simplified GET endpoint. Only supports URL-based submissions - no file uploads or base64 encoding. Faster and simpler than POST endpoint for basic use cases.
Input parameters
Publicly accessible URL of the image or PDF file to OCR. Must be a valid HTTP/HTTPS URL.
OCR language code: ara=Arabic, bul=Bulgarian, chs=Chinese Simplified, cht=Chinese Traditional, hrv=Croatian, cze=Czech, dan=Danish, dut=Dutch, eng=English, fin=Finnish, fre=French, ger=German, gre=Greek, hun=Hungarian, kor=Korean, ita=Italian, jpn=Japanese, pol=Polish, por=Portuguese, rus=Russian, slv=Slovenian, spa=Spanish, swe=Swedish, tha=Thai, tur=Turkish, ukr=Ukrainian, vnm=Vietnamese
If True, returns word-level bounding box coordinates (Left, Top, Height, Width) for each detected word. Useful for document layout analysis.
Output
Data from the action execution
Error if any occurred during the execution of the action
Whether or not the action execution was successful or not
Agents can call 3 OCR.space actions on Nagent: "Get Conversion Statistics", "Extract Text from Image/PDF (OCR)" and "Extract Text from Image URL (GET)". Get Conversion Statistics: Retrieve OCR API conversion statistics and usage data (PRO accounts only). Each action is listed on this page with its input parameters and its output.
OCR.space connects with an API key, under your workspace's own connection. Nothing is enabled on connect: each action is allowed one at a time and can be scoped to the agents that need it.
It also takes 1 optional input: startDate. It returns data, error and successful.