Capture, Selector, and projects
Visual sources, timeline, markup, connections, descriptions, history of changes and saving the project.
In the current source, General and GUI Automation support the main workflows. Industrial and UAV are present as experimental modes, but do not yet form complete production workflows. The current file metadata does not establish the capability set of a specific published installer.
| Mode | Status | In source | The main process is ready | Current boundary |
|---|---|---|---|---|
| General | Implemented | Yes | Yes | Markup, general computer vision techniques and project saving. |
| GUI Automation | Implemented | Yes | Yes | Preparation of interface data and separate execution of the tested script. |
| Industrial | Experimentally | Yes | No | Subject data types, method execution, validation, and confirmed markup changes are available. There is no completed production process. |
| UAV | Experimentally | Yes | No | You can work with frames, run methods, and check the results. Mission control, GIS and autopilot are not included in the mode. |
Visual sources, timeline, markup, connections, descriptions, history of changes and saving the project.
Classical methods, methods with models, sequences and graphs of methods, preview and saving of results. Some methods require separate models and dependencies.
Loading a project, searching for visual and textual states, waiting for conditions and acting through the selected input method. Running the script is always done separately.
The embedded code environment receives the selected task context. The agent can change files in the working directory, so changes must be checked and stored in a version control system or a separate copy.
The built-in assistant is disabled by default. Suggested changes are shown before implementation, and high-risk actions require confirmation.
Features and limitations →Computer vision techniques help determine the purpose of a recorded action. The proposed result is stored separately and requires human verification.
Local options and connections to external services with their own credentials are supported. Availability may vary by provider, network, and model selected.
Dictation, window bindings, text review, and optional screenshot processing with a language model are implemented. The Voice2Text page shows current download availability and file details.
Description Voice2Text →The account is used for profile, desktop program connection, support and available online operations. No login required to work locally.
Payment on the site is disabled. It is currently not possible to top up the balance displayed in your profile through the website.
YOLO, SAM, OmniParser and other model methods require suitable model files. Tesseract OCR requires local installation of Tesseract. A separate graphics card is not required for basic functions.
Cloud speech recognition and LLM require network and access to the selected service. Local connections are available with installed models.
Interface actions require a configured Arduino HID or FakerInput Virtual HID. Installing the FakerInput system driver is a separate action that requires Windows confirmation.