Skip to content

Transcription

Persyk uses speech-to-text AI to transcribe your voice recordings into text. After recording a voice memo, the audio is sent to a transcription service and the resulting text is copied to your clipboard.

Any OpenAI-compatible transcription endpoint can be used, giving you flexibility to choose between cloud services or self-hosted solutions based on your privacy, latency, and cost requirements.

Configuration

Configure transcription in your settings by specifying a provider and model:

{
"providers": {
"speaches": {
"type": "openai-compatible",
"baseUrl": "http://localhost:8000/v1",
"apiKey": "sk-..."
}
},
"transcription": {
"enabled": true,
"provider": "Systran/faster-distil-whisper-large-v3",
"model": "whisper-1"
}
}

Specify Language for Faster Transcription

Supplying a language request parameter can significantly improve transcription speed for both Speaches and OpenAI. When the language is specified, the model skips language detection and processes audio faster.

{
"transcription": {
"enabled": true,
"provider": "speaches",
"model": "Systran/faster-distil-whisper-large-v3",
"requestOptions": {
"language": "en"
}
}
}

Use ISO 639-1 language codes (e.g., en for English, es for Spanish, de for German).

Custom vocabulary

In Settings → Transcription → Custom vocabulary, enter a word or phrase and press Enter or click Add. Terms appear as compact chips; click a chip’s × to remove it. Spaces stay within a phrase (such as “Example phrase”). Paste a newline-separated list to add multiple terms at once; blank entries and exact duplicates are skipped. Unfinished input is added when you leave the input field.

You can also set transcription.keywords in settings.json5:

{
transcription: {
provider: "openai", // Must match a configured provider
model: "gpt-transcribe",
keywords: ["Persyk", "SvelteKit", "voice memo"],
},
}

OpenAI documents keywords support for gpt-transcribe and gpt-live-transcribe, not Persyk’s default gpt-4o-mini-transcribe. For other OpenAI-compatible providers, check their model support. Persyk forwards the list without converting it to a prompt or changing your model.

Vocabulary guides recognition; it does not guarantee exact spelling. Changes apply to the next recording. Remove the setting or use keywords: [] to stop sending vocabulary hints.

Provider Options

Speaches is another project of mine - an OpenAI-compatible inference server with support for transcription, translation, text-to-speech, voice activity detection (VAD), speaker embedding, Realtime API, and more. See the GitHub repository for setup instructions.

Best for users who want full control over their data and are willing to run a local server. With the right hardware, you can achieve a significantly lower latency than cloud alternatives.

Hardware Recommendations

For low latency, having a GPU is recommended, but CPU is also supported.

CPU users: To optimize for latency over accuracy, use smaller models:

  • Systran/faster-whisper-tiny.en
  • Systran/faster-whisper-tiny
  • Systran/faster-whisper-small.en

GPU users: Larger models provide better accuracy while maintaining good latency:

  • Systran/faster-distil-whisper-large-v3
  • Systran/faster-whisper-large-v3
  • deepdml/faster-whisper-large-v3-turbo-ct2

Example configuration:

{
"providers": {
"speaches": {
"type": "openai-compatible",
"baseUrl": "http://localhost:8000/v1",
"apiKey": null,
},
},
"transcription": {
"enabled": true,
"provider": "speaches",
"model": "Systran/faster-whisper-small.en",
},
}

OpenAI

Use OpenAI’s API directly. The quickest way to get started - just add your API key and you’re ready to go.

Example configuration:

{
"providers": {
"openai": {
"type": "openai-compatible",
"baseUrl": "https://api.openai.com/v1",
"apiKey": "sk-...",
},
},
"transcription": {
"enabled": true,
"provider": "openai",
"model": "gpt-4o-mini-transcribe", // or "whisper-1"
},
}