feat: audio transcribing

This commit is contained in:
2026-07-15 14:45:41 +07:00
parent 4e6af66362
commit 78f31e4a65
3 changed files with 104 additions and 9 deletions

View File

@@ -1,6 +1,6 @@
# Telegram OpenAI-Compatible Bot
Minimal aiogram bot that forwards whitelisted Telegram users' text messages to an OpenAI-compatible chat completions API.
Minimal aiogram bot that forwards whitelisted Telegram users' text and voice messages to an OpenAI-compatible API.
## Features
@@ -11,6 +11,7 @@ Minimal aiogram bot that forwards whitelisted Telegram users' text messages to a
- Persistent per-user context in SQLite
- `/new` command to reset the current user's context
- Per-user message batching before sending to the model
- Voice message transcription through a separately configured transcription model
- Model-controlled multi-message replies with `<break>` tags
- Configurable model and system prompt
- API timeouts, retries, response limits, and safe user-facing fallbacks
@@ -37,6 +38,7 @@ BOT_TOKEN=123456:telegram-token
API_BASE_URL=https://api.openai.com/v1
API_TOKEN=sk-your-api-token
MODEL=gpt-4o-mini
TRANSCRIPTION_MODEL=whisper-1
WHITELIST_USER_IDS=123456789
DATABASE_PATH=data/bot.sqlite3
MAX_CONTEXT_MESSAGES=20
@@ -45,6 +47,8 @@ MESSAGE_BATCH_DELAY_SECONDS=3
`API_BASE_URL` must be the API root. Do not include `/chat/completions`; the OpenAI SDK appends that path automatically.
`TRANSCRIPTION_MODEL` is used for Telegram voice message transcription before the transcript is sent to `MODEL`.
For OpenRouter, use:
```dotenv
@@ -75,6 +79,12 @@ The bot stores successful user/assistant turns in SQLite at `DATABASE_PATH`. Eac
Send `/new` to delete your stored messages and start a fresh conversation.
## Voice Messages
Telegram voice notes are downloaded, transcribed with `TRANSCRIPTION_MODEL`, then the transcript is passed through the same batching and chat-completions flow as text messages.
Telegram does not expose a bot API to mark a voice note as listened in the user interface, so that step is not supported.
## Message Batching
`MESSAGE_BATCH_DELAY_SECONDS` delays the API request after each incoming Telegram message. If the same user sends more messages during that delay, the timer restarts and the messages are sent to the model together as separate entries: