Recited ayat are matched, not just transcribed
The hard part of an Islamic video clipper is not the clipping. Whisper hears sound, and for recitation that is not enough — a phonetic transcript of an ayah is not the ayah. Recited Quran is matched against a corpus of all 6,236 ayat, and the clip carries the verse itself, set in the Amiri typeface, with its translation on the line beneath. Captions are burned in with libass and timed to the words as spoken, so the text keeps pace with the reciter. The Quran Recitation template goes further and captions scripture and nothing else, so no half-heard aside sits under a verse.
An English talk with Arabic inside it is two languages
Most transcription picks one language from the opening seconds and applies it to the whole file. An English lecture with two minutes of recitation in the middle comes back as Latin nonsense, and after that nothing can tell it was ever Arabic. Auto-detect here is multilingual and switches per segment. Arabic speech that is not scripture is captioned in Arabic with an English line beneath it, taken from a second translate pass. You can also pin the language — English, Arabic or Urdu — when you already know what is on the recording.
Scripture cannot leave without a person seeing it
Every clip lands in a review queue and waits for a human decision. Nothing publishes on its own. Clips containing scripture carry an extra flag, QUOTE_RISK, that forces review and never bypasses — not even under automatic approval. The reviewer watches the rendered clip, the same file that would post, with its score and the model's short reasons beside it: complete ending, question hook, stands alone. Approve, reject or send it back, with keyboard shortcuts for a long queue. A clip that needs a trim or a section cut out of the middle opens in the editor from the same queue, and Save renders the cut.
What happens between the link and the first clip
Give it a YouTube URL or an MP4, then pick the start and end of the stretch worth clipping — thirty seconds minimum. Only that stretch is downloaded and processed, so a three-minute selection from a ninety-minute lecture costs three minutes. Whisper transcribes it with word-level timings. A self-hosted Ollama model scores candidate moments and returns its reasons, and cuts land on complete moments rather than fixed intervals. Face detection crops 16:9 to vertical 9:16, a nasheed is mixed under the speech and ducked beneath it, and captions burn in word-by-word, karaoke, phrase or stacked, across five templates.
Approved clips go out on your own channels
An approved clip is scheduled into posting windows — four a day, eight on Studio — and published to the YouTube, TikTok, Instagram and Facebook accounts you have connected. Each destination reports its own state, so a clip that went live on YouTube is not filed as a failure because TikTok refused it. What you will not find here are view counts or watch time: no connected platform sends that data back, and there is no dashboard here inventing it.
Questions
How much of my lecture am I charged for?
One token is one source minute, and you are charged for the stretch you selected, not the length of the video it came from. Basic is free — forty tokens over a seven-day trial. Re-rendering a clip and cutting more clips from the same lecture never cost tokens again. A render that fails is not charged, and neither is a clip you reject before it exports.
What happens to a clip that contains a Quran verse?
The recitation is matched against the 6,236-ayah corpus and rendered as the verse in Amiri with its translation beneath. The clip is then flagged as carrying scripture and forced into human review, which no automation setting can switch off. If you use the Quran Recitation template, scripture is the only thing captioned — the surrounding lecture is left off the screen.
Can I turn the background nasheed off?
Yes, but you have to say so. A nasheed is mixed in by default and ducked under the speech, so the words stay clear; music is on unless it is explicitly switched off for that job. The rest of the look is set by the template you choose — Bold Stack, Clean Line, Headline, Mono Minimal or Quran Recitation — and the caption mode you pick with it.
Why not just use a general AI clipper?
Three things. Recited Quran is matched against a 6,236-ayah corpus and set as the verse with its translation, rather than transcribed phonetically. Arabic speech that is not scripture is captioned in Arabic with an English line under it. And any clip containing scripture is forced into human review, whatever the automation settings say. The clipping itself — transcription, scoring, framing, captions — is the ordinary part.
Basic includes the whole workflow.
Import a source, generate clips, review every one and publish to your own connected channels. Upgrade when you need more.

