GiliSoft AI Toolkit User Guide

AI Toolkit User Guide for Windows

Choose the right AI workspace, add source material, review the result, and save accepted output.

  • Extract and transform text and speech
  • Create authorized voice and face projects
  • Clean, cut out, repair, and enhance images
GiliSoft AI Toolkit for OCR, speech, voice, face, and image tasks

Complete an AI task in four steps

Start with one representative file or short sample. Check the result before processing important or larger source material.

Choose a workspace

Open the dashboard and select the tool that matches the required output.

Add the source

Load text, audio, an image, or another supported source and confirm it is correct.

Set the options

Choose language, voice, selection, quality, or processing settings for the task.

Review and save

Inspect the generated result, correct visible problems, and export it to a new file.

Choose the AI Toolkit workspace

Each workspace has a separate input and output. Use the task name below to open the correct tool instead of sending every source through the same process.

AI Toolkit OCR workspace

OCR

Extract editable text from screenshots, scans, photographed pages, and image-based documents.

AI Toolkit text to speech workspace

Text to Speech

Turn written scripts into spoken audio for drafts, lessons, accessibility, and narration.

AI Toolkit speech to text workspace

Speech to Text

Transcribe supported recordings into editable text for notes, interviews, and documentation.

AI Toolkit voice cloning workspace

Voice Clone

Create a custom voice only from recordings you own or have permission to use.

AI Toolkit face swap workspace

Face Swap

Prepare authorized photo and video face projects and inspect identity and edge consistency.

AI Toolkit background cutout workspace

Background Cutout

Separate a subject from its background for product images, profiles, and design assets.

AI Toolkit image cleanup workspace

Image Cleanup

Remove unwanted marks or objects from images you are authorized to edit.

AI Toolkit dashboard

More AI Tools

Return to the dashboard for photo repair, enhancement, subtitle, and other available utilities.

Create text and voice output

Match the source, language, and output

  1. Open OCR, Text to Speech, Speech to Text, or Voice Clone from the dashboard.
  2. Add a clear source image, script, or audio sample.
  3. Select the correct language and available voice or recognition options.
  4. Process a short sample and check names, numbers, punctuation, pronunciation, and timing.
  5. Edit the source or settings when the sample needs correction.
  6. Export the approved text or audio to a clearly named file.
AI output can contain recognition or pronunciation errors. Review important names, dates, figures, and quoted material before use.
Create speech from text in GiliSoft AI Toolkit

Prepare face and image results

Use clear source material and inspect the boundaries

  1. Open Face Swap, Background Cutout, Image Cleanup, Repair, or Enhancement.
  2. Load an image or video you own or are authorized to edit.
  3. Choose a source with clear subjects, useful resolution, and suitable lighting.
  4. Mark the required subject or cleanup area where the tool provides a selection control.
  5. Generate a preview and inspect hair, hands, face edges, text, reflections, and repeated textures.
  6. Save the result as a new file so the original remains available.
Do not use face or voice tools to impersonate another person or misrepresent consent. Keep the source and generated output identifiable in sensitive workflows.
Remove an image background with GiliSoft AI Toolkit

Review AI output before delivery

Accuracy

Check text, speech, faces, object edges, colors, and important visual details.

Permission

Confirm you may use the source voice, face, image, audio, or document.

Originals

Keep the unmodified source until the generated result has been accepted.

Destination

Open the exported file in the application or device where it will be used.

AI Toolkit troubleshooting and FAQ

Why is OCR text inaccurate?

Use a sharper, straighter image with readable contrast, select the correct language, and review complex layouts manually.

Why does generated speech pronounce a word incorrectly?

Check the selected language and voice, adjust spelling or punctuation, and test a shorter sentence before exporting the full script.

Why does a cutout have rough edges?

Start from a larger image with clearer subject separation and inspect hair, transparent areas, and shadows at full size.

Why does a face result look unnatural?

Use sources with compatible angle, lighting, expression, and resolution, then inspect several frames when processing video.

Should I overwrite the original file?

No. Save generated output separately until it has been reviewed and accepted.

Ready to try an AI workflow?

Choose the matching workspace, add a representative source, review the generated result, and save the accepted output separately.