Centurion

CENTURION

SPEECH IN · SUBTITLES OUT · .NET 10

Features

Subtitle Styles (Batteries Included)

The default ASS template ships with two ready-made styles, tuned for dual-language subtitles:

Style Role Font Size Position
Default Main line (source language) Arial (bold) 84 Top-ish, MarginV 100
Sub Secondary line (translation) Microsoft YaHei 81 Bottom, MarginV 28

Speaker Diarization (Who’s Talking?)

Every pipeline path can label who said what — each word gets a Speaker attribute.

Two backends, both implemented natively through the CrispASR CLI (no Python required ):

Backend DiarizationBackend Method Notes
CrispASR built-in crispasr foxnose (default), energy, xcorr, vad-turns Zero extra deps, auto speaker count
Pyannote + TitaNet pyannote pyannote segmentation + TitaNet embeddings Rock-solid on long audio; models auto-downloaded

Configure via WorkflowConfig:

DiarizationBackend = "crispasr",   // "none" to disable
DiarizationMethod  = "foxnose",    // crispasr methods
DiarizationModel   = "pyannote-seg-3.0",
NumSpeakers        = 0,            // 0 = auto

Diarization runs before sentence splitting and is non-fatal: if it fails, subtitles still get generated (just without speaker labels). No drama.


Vocal Separation (BGM, Step Aside!)

Background music drowning out the speech? Separate the vocals first, then transcribe — clean input, better subtitles.

Centurion asr song.mp4 --vocal-separation
Centurion asr song.mp4 --vocal-separation --vocal-separation-model htdemucs_ft

GPU Detection & Auto-Download

Centurion detects your hardware at startup and downloads the right tool builds automatically. No more “please manually install the CUDA version” emails.

Centurion asr audio.mp3 --device cuda     # force CUDA builds
Centurion asr audio.mp3 --device cpu      # stay cozy on CPU

Startup banner example:

 Device: Platform: win-x64; GPU: Intel(R) Arc(TM) 130T GPU (16GB) (non-NVIDIA); RAM: 15.7 GB; Recommended: Vulkan

Metadata Registry

Tools & models live in an external JSON registry, loaded at startup — edit it without touching code.

{
  "tools": {
    "whispercpp": {
      "downloadUrl": ".../whisper-bin-x64.zip",
      "variants": {
        "cuda": { "downloadUrl": ".../whisper-cublas-12.4.0-bin-x64.zip" }
      }
    }
  }
}

Non-Latin Language Support

Centurion no longer assumes your audio speaks English with spaces.

Centurion asr 讲座.wav -l zh --transcriber qwen3-asr-1.7b
Centurion asr anime.mkv -l ja

Architecture

Operator-pipeline architecture: every stage is an independent module — swappable, extendable, reorderable. No monoliths, no tears.

 input ──► FFmpegConvert ──► AudioPreprocess ──► [ VocalSeparation] ──► Transcribe
              ──► [ Diarization] ──► SentenceSplit ──► TextCleaning ──► Alignment ──► ASS

Project Layout (clean layering, zero circular deps):

Project Role Depends on
Centurion.Models Pure data models (Sentence/Word/WorkflowConfig/ASS…), metadata JSON registry, console facade, InferenceDevice nothing
Centurion.Abstractions Interfaces & abstract bases (operators, strategies, factories), exceptions, request DTOs Centurion.Models
Centurion.Core The engine: managers, operators, pipeline executor, strategies, DI wiring Centurion.Models + Centurion.Abstractions
Centurion.Cli Spectre.Console command-line front-end Core + Models + Abstractions
Centurion.Server ASP.NET Core REST API over the packaged commands Cli + Core + Abstractions
Centurion.Tests xUnit test suite Core + Models + Abstractions

Notes & Limitations