PixLab Vision Platform AI Agents & Desktop Automation

Controlled desktop AI agent workflows with Think-Act

Think-Act turns a defined desktop task into a controlled agentic workflow. Give it the goal, relevant context, and optional success criteria. It works with the skills and tools available to the run, pauses when configured actions need approval, and leaves activity and results for you to review.

Think-Act product screenshot

Placeholder for a future Think-Act workspace product screenshot.
What makes work agentic

More than a response. A controlled path through the work

A normal AI response may finish when text is generated. Agentic work has an objective to complete, context to understand, capabilities to use, boundaries to respect, and evidence you can inspect before deciding the task is done.

Agentic does not need to mean uncontrolled autonomy. The useful pattern here is bounded work + explicit access + human review.

AI response

Prompt in. Response out.

Useful for answering, drafting, rewriting, summarizing, or explaining something when no further action is required.

Receives a question or prompt.

Uses the context supplied to it.

Returns a result for you to use or edit.

Think-Act workflow

A task with a finish line.

Think-Act treats agent work as a controlled process with a defined objective and available capabilities, not an unlimited instruction to act.

Task

Defined outcome

Success

Optional criteria

Context

Real source material

Access

Enabled tools only

Policy

Allow, ask, block

Evidence

Activity + result

The Think-Act workflow

From a bounded request to a result you can check

The workflow starts with the work itself. Think-Act uses the context and capabilities available to the run, surfaces limitations when something is missing, and keeps important decisions with the user.

Agent workflow screenshot

Placeholder for a future Agent workflow screenshot.
  1. 01

    Define the work

    Provide a bounded task

    Describe the result you want, what the agent should work with, and any boundaries it should follow. Keep the scope narrow enough that the result can be checked.

  2. 02

    Bring context

    Give it the material the task depends on

    Selected text, files, PDFs, scans, screenshots, images, recordings, and other supplied material can become working context.

    Text Files Documents Screenshots Recordings
  3. 03

    Understand the requirement

    Determine what the task actually needs

    Think-Act works out which parts of the task can be handled with the context, built-in capabilities, reusable skills, or enabled tools available to the run.

  4. 04

    Use capabilities

    Use only available skills and enabled tools

    Skills can guide the method. Tools provide real capabilities. A workflow cannot perform an external action through a skill alone when the required tool is unavailable.

  5. Human decision

    Pause when an action requires approval

    Configured consequential actions can wait for confirmation before they execute. Security policy determines whether a supported action is allowed, asks first, or is blocked.

  6. Inspect the run

    Review activity, observations, and the result

    Tool observations, supported activity events, and results remain available to inspect after the run. Important output should be checked against its source and success criteria before use.

Capabilities for desktop work

Use each capability on its own or as context for an agent task

Think-Act combines a Desktop Agent with focused tools for writing, OCR, document intelligence, and screen recording. Use the focused tool when that is enough, or carry its output into a broader agent task.

Multi-step work

Desktop Agent

Plan and review a bounded desktop task

Provide the request, source material and limits. Think-Act can coordinate the capabilities available to the run, expose supported activity for review and pause where the configured policy requires confirmation.

  • Optional success criteria for a clear finish line
  • Visible supported observations and execution activity
  • Approval controls for configured consequential actions
  • Reusable local skills after the process is reviewed
Explore the Think-Act Desktop Agent

Write, rewrite, translate or summarize

Work from selected or supplied text, then compare and edit the resulting draft before using it.

Explore writing assistance

Extract and review document content

Use OCR with supported scans, screenshots and image-based PDFs, then organize checked content into fields, tables or summaries.

Explore document intelligence

Record a process and prepare documentation

Capture a workflow, then prepare a transcript, SOP, checklist or bug report that can be compared with the recording.

Explore screen recording
Skills, tools & connectors

Instructions are not capabilities. Each layer has a different job

Skills provide reusable methods. Tools provide real capabilities. Connectors expose those tools to the agent. Policy then decides whether a supported action can proceed, must ask first, or stays blocked.

Layer 1 · Instructions

Skills guide the method

Local instructions can describe the sequence, checks, prerequisites and expected format for familiar work.

  • Shape how the workflow is approached
  • Preserve reviewed checks and output requirements
  • Cannot perform an external action alone

Layer 2 · Capabilities

Tools do the work

A tool is the callable capability that can read, create, change, query, or otherwise interact with an allowed resource.

  • Available tools define the real action boundary
  • Another application needs its supported tool
  • A skill cannot substitute for a missing tool

Layer 3 · Local access

Local MCP connectors expose tools

A local MCP server can expose tools to Think-Act while it is installed, trusted, enabled, connected, and healthy. Local packages run code on your device and require your trust.

  • Tools appear only while the connector is available
  • Unavailable or failed tools are not exposed
  • Missing capabilities are reported instead of invented
Security and human control

Know what is available, what needs approval, and what happened

Think-Act limits automation through explicit tool availability, configurable action policies, and reviewable activity. The agent cannot use a tool you have not installed, trusted, and enabled.

Enable only what the task needs

Connectors unavailable to a run cannot be used by it.

Allow, ask, or block per action

Set each supported action type according to its impact.

Review what the agent did

Inspect tool calls, results, and observations after a run.

Unavailable connectors cannot be used

They stay excluded until the connection is restored and re-enabled.

Approval Controls Screenshot

Placeholder for a future Approval Controls Screenshot.

Recommended starting policy

Allow

Read selected content

Ask

Create or change files

Ask

Send or publish

Block

Delete or risky commands

Practical agent workflows

Choose work with a source, a boundary, and a checkable result

The best starting points are narrow enough to review without guesswork. Keep external writes out of the first run, then add a trusted tool only when the workflow genuinely needs it.

No connector needed

Rewrite a weekly update

Provide the draft and the tone or format you want. Review the rewrite against the original before sending it.

Rewrite with Think-Act
No connector needed

Summarize a report or transcript

Attach the source and choose the length and focus. Check the key points before sharing the summary.

Summarize with Think-Act
No connector needed

Extract fields from a receipt or invoice

Attach an image or PDF and list the fields you need. Compare the result with the original document.

Extract receipt and invoice fields
No connector needed

Review a scanned contract

Provide the scan and the clauses or terms to check. Review the summary before making a decision.

Review scanned documents
No connector needed

Record a process and prepare an SOP

Record the workflow, with optional narration, then use it to draft a step-by-step SOP or checklist.

Create an SOP from a recording
Connector required

Organize a small folder with clear rules

Define the sorting or naming rules. An appropriate supported connector is required before Think-Act can move or rename files.

Tools and connectors
Illustration of PDF, Word, presentation, Excel, and HTML files being converted into structured, LLM-ready text and JSON output for agentic applications
PixLab's document APIs prepare varied formats for RAG, extraction, and agentic tasks — independently of Think-Act.
For developers & builders

Document intelligence for agentic applications

PixLab’s LLM and document services can transform source documents into structured material for search, extraction, RAG preparation and downstream application workflows.

OCR Document parsing Segmentation Layout analysis Structured extraction RAG preparation Chunking Embeddings

Three distinct products, one platform. Think-Act is the desktop agent. PixLab APIs are developer building blocks. Vision Workspace provides browser-based document tools. They share document intelligence capabilities but are used independently.

For supported identity-document extraction, review the separate DocScan API.

FAQ

Questions about Think-Act and agentic workflows

Clear answers about bounded tasks, human control, skills, tools, connectors, and developer building blocks. Can't find the answer you need? Contact support.

What is Think-Act?

Think-Act is PixLab's Windows desktop AI workspace for writing assistance, OCR and document review, screen recording, reusable local skills, and bounded multi-step tasks. It is designed to keep work visible and reviewable — not to operate as an uncontrolled autonomous agent.

Is Think-Act an autonomous agent?

No. Think-Act is designed around bounded, human-controlled workflows. You define the task, choose which tools are available, and configure whether consequential actions require your confirmation. The agent reports what it did and keeps activity available for review.

What is a bounded agent task?

A bounded task has a clear goal, optional success criteria describing what done looks like, defined source material, and explicit limits on what the agent may change or access. Think-Act works through the task using only enabled tools and keeps the activity visible so you can verify the result.

Does Think-Act need connectors to work with applications?

Not for every task. Writing assistance, OCR, document review, screen recording, and local skills can work without external connectors. Actions inside another application or external system require the appropriate supported connector to be installed, trusted, enabled, connected, and healthy. If it is unavailable, Think-Act reports the limitation instead of inventing the outcome.

What is the difference between a skill, a tool, and a connector?

A skill provides reusable instructions for how to approach a task. A tool provides an actual capability. A connector exposes tools to Think-Act. Skills cannot bypass missing tools, and local MCP connectors can run code on your device, so install and enable only packages you trust.

How are consequential actions handled?

Security policy determines whether an action may proceed automatically (Allow), must pause for your confirmation (Ask), or is blocked entirely (Block). Write, delete, and network actions can be placed behind an Ask gate so the workflow waits for your decision before proceeding. Review Think-Act's permission model for the full detail.

Can developers build agents with PixLab APIs?

Yes. PixLab provides APIs for LLM-ready document parsing, OCR, document segmentation, chunking, embeddings, structured extraction, and vision-language analysis. These can be used independently of Think-Act to build custom agentic applications. Start with the VLM API reference or the LLM and parsing APIs.

Your next step

Start with Think-Act and keep the work reviewable

Download Think-Act for controlled desktop work. Use the supporting links to inspect features and security first, or move to PixLab's APIs when the workflow belongs inside your application.