Subtitle tools for agents · Developer preview

Subtitle MCP Server — OCR, Transcribe, Translate and QC Video Subtitles with Claude/Codex

By Flora Wang, GeekLink team · Updated October 1, 2026

Give your AI assistant a subtitle workflow, not another dashboard to navigate. GeekLink’s local MCP preview lets an agent extract burned-in text, transcribe speech, translate subtitles, inspect suspicious cues and export the result.

Availability: developer preview, not a public release. The local server is configured on our development Mac for Claude Desktop and Codex. It is not bundled in the public GeekLink installer, and no public remote endpoint is available. The examples below describe the current implementation, not an installation you can complete today.

Contact us about early access

Tell us your assistant, operating system and the subtitle workflow you want to automate. No purchase required to register interest; access and timing are not guaranteed.

For a quick overview and early-access details, visit the GeekLink Subtitle MCP product page. This guide covers the current tools and workflow constraints in more detail.

TL;DR: A subtitle MCP server exposes subtitle operations as tools an AI assistant can call. GeekLink’s developer preview connects an assistant to the running desktop app: import videos, start OCR or speech tasks, read subtitle results, inspect video frames, edit through existing controls and export. Its most useful loop is extraction followed by targeted review against the original picture. Public installation is still pending.

What is a subtitle MCP server?

MCP means Model Context Protocol. For a subtitle app, the practical idea is simple: an assistant calls a defined tool and receives results it can use for the next step. It does not need to guess where an “Extract” button is on every screen.

The assistant plans the workflow; GeekLink performs subtitle processing. The preview runs alongside the desktop app and connects to a local MCP client over STDIO. Task status, subtitle text and requested video-frame images return to the assistant for review.

This matters because “get the subtitles” can mean different things. Text burned into a picture needs video subtitle OCR. Spoken dialogue needs transcription. Existing subtitle text needs translation or editing. An agent should choose the right operation rather than treat all video inputs as the same problem.

Which atomic subtitle tools can an agent call?

These are the names in the current developer preview. OCR, transcription and translation share a task launcher; they are not separate public tools named extract_hardcoded_subtitles or transcribe_video.

OperationCurrent toolWhat the agent gets
Discover capabilitiesget_capabilitiesAvailable actions and their input fields.
Import and list videosimport_videos
list_videos
The app’s video library and local file information.
OCR, transcribe, translatestart_taskA task ID for ocr, speech, translate or speech_translate.
Track or stop processingget_task
stop_task
Progress, status and the next OCR review step.
Review OCR regionsget_ocr_line_candidates
confirm_ocr_lines
Candidate images, per-line languages and explicit keep/delete decisions.
Read and verify subtitlesread_subtitle
inspect_video_frame
Source text, translations, review flags and a frame at a requested time.
Edit in the appget_ui_state
ui_action
run_action
Existing controls and supported app actions, rather than a second editing system.
Export text or videoexport_subtitles
get_export_status
SRT, TXT, burned-in video or toggleable subtitle tracks; video-export progress.

The important capabilities beyond extraction are visibility and control: inspect a frame, read only flagged cues, find out whether a task actually finished, and stop it when needed. Those small operations make a reviewable workflow possible.

How can Claude or Codex extract and review burned-in subtitles?

Burned-in subtitles are text inside video pixels, not a track that can simply be copied. The preview exposes the subtitle extraction and review workflow, including the choice of which screen text should count as subtitles.

  1. Check the library. Import the intended local files and inspect the library before starting a batch.
  2. Start OCR with an explicit source language. Poll the task until the candidate-review step is ready.
  3. Look at the candidate images. Inspect the supplied frame and line crops. Distinguish dialogue from logos, character names or menus.
  4. Confirm the regions. Assign a language to each retained line. Give a reason for every rejected candidate. Keep at least one subtitle line or add a suitable manual region.
  5. Review the extracted cues. Read flagged subtitles and inspect the corresponding original frames. A suspicious cue is a reason to look, not permission to invent missing text.
  6. Save and export. Save any editor changes, choose SRT or a video export, and verify completion before reporting success.

Example request after a preview connection is configured

“Check the imported Japanese clips. Extract the burned-in dialogue, ignore channel logos, and inspect every flagged cue against its video frame. Show me uncertain readings before changing them. Export a source-language SRT after I approve the review.”

This is an example of intended orchestration, not a benchmark or a guarantee of unattended accuracy. The agent must inspect the evidence and the actual task results.

Can the same MCP server transcribe and translate video subtitles?

Yes, in the developer preview. Use a speech task when the source is spoken audio, a translation task when source subtitles already exist, or a combined speech-and-translation task. The app retains its existing language, model and quota checks.

Local video transcription produces timed source subtitles. Cloud subtitle translation works from subtitle text using the chosen service. The assistant should ask for the target language and the desired source, target or bilingual output instead of silently assuming English.

Choose OCR when you need the words actually shown on screen; choose speech recognition when you need what was said. They can disagree, especially when a video already contains shortened or translated captions. Combining them without deciding which source is authoritative can create a convincing but incorrect subtitle file.

What does subtitle QC mean for an AI agent?

QC means quality control before publication. A completed processing task does not mean every subtitle is correct. The agent can read source and translated cues, narrow its attention to review flags, and request a video frame when the evidence is visual.

For OCR, check unclear characters, missing lines and unrelated on-screen text. For translation, compare names, numbers, negation and meaning with the source. GeekLink’s post-translation QC workflow explains why a valid-looking subtitle file can still contain omissions.

The preview does not expose a separate “perfect subtitles” tool. read_subtitle can filter review flags; frame inspection and the existing editor support the follow-up. Flags are not exhaustive, and a frame alone cannot verify what an unseen speaker said. Ask the agent to leave uncertainty visible and explain proposed corrections.

Before burned-in export, check the result yourself: once text is rendered into the picture, viewers cannot switch it off. Current MCP burned-in export uses fixed export styling; do not assume every preview-style adjustment changes the rendered video.

How do Claude, Codex and ChatGPT connect?

Claude Desktop and local Codex: our development setup registers the local STDIO server with each client. GeekLink must be running on that computer. Public setup instructions and an installer entry point are still pending, so we are not publishing a copy-and-paste installation command for a package you cannot download.

ChatGPT on the web: this preview does not offer a public remote MCP service. Configuring a local Codex client does not automatically expose the app to ChatGPT web. A remote connection is a separate integration, not something this page enables.

Codex’s own MCP connection documentation describes its local STDIO and HTTP options. That explains the client mechanism; it does not mean the GeekLink preview is publicly installable.

What should you know before requesting early access?

  • Development Mac first. The preview is used on our development Mac. A user-ready Mac package and Windows MCP launcher are not shipped yet.
  • Paid access gate, not a launch offer. The current implementation requires paid Pro or a valid Local Pass. Trial access is not sufficient. Do not buy a plan solely to obtain a feature that has not been publicly released.
  • Batch scope matters. start_task processes library videos still missing the requested results; it is not a single-video selector. Check the library before starting. Only one heavy processing task runs at a time.
  • Local processing is not an offline-assistant promise. OCR and transcription run on the computer. Subtitle text and requested frame images are supplied to the connected assistant; its own data handling applies. Cloud translation also sends subtitle text to the selected service.
  • Separate access and usage costs. A Local Pass does not include cloud AI translation credits. MCP does not bypass model availability, translation credits or normal export checks.
  • Confirm outcomes. A sent click is not proof of a saved edit. Save changes before export and check task or export status before calling the workflow complete.

For a workflow you can use from the public app today, see Watch Folder automation for speech and translation. That is a separate, folder-based workflow, not the MCP preview.

Subtitle MCP Server FAQ

Can I install the GeekLink Subtitle MCP Server today?

Not from the public installer. The MCP server is currently a developer preview used on our development Mac, with local configurations for Claude Desktop and Codex. Contact the team to discuss early access; there is no public one-click MCP installation yet.

Can ChatGPT on the web control my local GeekLink app?

Not through this local preview. A local Codex MCP connection is not a ChatGPT web connection. GeekLink does not currently provide a public remote MCP endpoint for ChatGPT web.

Does an agent need screenshots to extract burned-in subtitles?

Yes. In the OCR workflow, the agent inspects candidate subtitle images, chooses which lines to keep, assigns their languages, and explains any rejected lines before confirming extraction. It can inspect video frames again when reviewing suspicious cues.

Is subtitle QC automatic correction?

No. Review flags identify cues worth checking; they are not proof of an error. An agent can compare source text, translations and relevant frames, then propose or apply edits through the existing editor. Save changes and review uncertain cases before exporting.

Is the MCP preview free, and does it keep everything local?

The current preview requires paid GeekLink Pro or a valid Local Pass; a trial does not unlock MCP. A Local Pass does not include cloud AI translation credits. OCR and transcription run on the computer, but tool text and requested frame images are shared with the connected assistant. Cloud translation sends subtitle text to the selected service. Do not purchase a plan solely for an unreleased MCP integration.

Disclosure: GeekLink is our own desktop app. This page documents our local MCP developer preview as of September 30, 2026, not a public release or a guarantee of subtitle accuracy.

What would you ask your subtitle agent to do?

Share your OCR, transcription, translation or review workflow. We are using concrete requests to shape the first public MCP release.

Discuss an agent workflow

Prefer the desktop workflow now? Explore the public GeekLink app. Its public installer does not include MCP.