How Clips Are Chosen

Highlights picked by what happens, not by loudness

Loudness meters clip loading screens and meaningless screams. Jitso Studio reads the speech transcript and looks at the footage with a local vision model, guided by a per-game judge written in plain English.

🗣

1. Speech Transcript

Local speech recognition transcribes your call-outs and reactions, and a language model scores the moments a viewer would enjoy.

👁

2. Vision Judge

A local vision model looks at frames from each candidate to check it really is a kill, a win or a clutch, and not a menu or a loading screen.

🎯

3. Per-Game Judges

60 games ship with a judge: plain-English notes on what a win, a big play and dead time look like in that game. Edit them or write your own. Hunt: Showdown also uses an on-screen kill-marker detector.

60 game judges ship with the app. Browse them or create your own.
Step by step

From recording to reviewed clips

Nothing is rendered until you approve it.

Step 1

Feed it footage

A local recording from your stream folders, or the link to one of your own YouTube VODs, fetched with yt-dlp.

Step 2

Analyze

Speech is transcribed locally with faster-whisper. A language model scores the moments, then a vision model checks frames from each candidate.

Step 3

Review

A list of candidates with scores and descriptions. Preview each one, Keep or Delete it, move In and Out, or Drop a stretch inside it.

Step 4

Generate

Approved clips render as 1080×1920 Shorts, or feed the Stitcher, Long Studio and Upload Studio.

The engine

Built-in local AI, or your own models

The default needs no account, no key and no credits. Everything else is optional.

Built-in llama.cpp

The default engine runs as a background service on your PC. It offloads the model to your GPU and frees that memory before heavy render and transcribe jobs.

Hardware-aware model catalog

The app detects your GPU and VRAM and recommends a model size (Gemma, Qwen or Llama). Models download from Hugging Face with their vision projector paired automatically.

Bring your own

Use Ollama, LM Studio or any OpenAI-compatible server, or OpenAI, Google Gemini or Anthropic with your own key. Model lists are fetched live from the provider.

Speech stays local

Transcription always runs on your PC with faster-whisper (CrisperWhisper is supported as an alternative). Only the model questions can move to a cloud provider, and only if you choose one.

Kill-marker detection

For Hunt: Showdown, a template-matching detector reads on-screen kill and hit markers from the video to timestamp eliminations exactly, alongside its judge.

Write your own judge

A judge is plain-English notes on what a win, a big play and dead time look like in a game. Edit the 60 that ship with the app or create a new one.

Transform your creator workflow today

Stream live, let AI curate your highlights, and publish your clips to YouTube and other platforms.

Get Started