---
title: "Speech to Text"
description: "Transcribe audio or video with precise word-level timestamps"
updated: 2026-09-10
---

# Speech to Text

**Node type:** `audio:whisper`  
**Category:** `audio`

## Description

Transcribe audio or video with precise word-level timestamps

## Pricing

- **Cost:** this node itself costs no credits

## Canvas ports

These appear as port handles on the left side of the node.

| ID          | Label             | Details              |
| ----------- | ----------------- | -------------------- |
| `audio_url` | **Audio / Video** | `VIDEO` _(required)_ |

## Sidebar config

These render as form fields in the right-side config panel when the node is selected.

| ID            | Label           | Details                                           |
| ------------- | --------------- | ------------------------------------------------- |
| `task`        | **Task**        | `TEXT` · options: `transcribe`, `translate`       |
| `language`    | **Language**    | `TEXT` · options: ``, `en`, `es`, `fr`, `de`, +95 |
| `chunk_level` | **Chunk Level** | `TEXT` · options: `none`, `segment`, `word`       |

## Outputs

| ID                | Label               | Type    |
| ----------------- | ------------------- | ------- |
| `text`            | **Transcript**      | `TEXT`  |
| `chunks`          | **Timed Chunks**    | `ARRAY` |
| `word_timestamps` | **Word Timestamps** | `JSON`  |

---

_Auto-generated from the Wireflow node registry._

---

Documentation index: fetch https://www.wireflow.ai/llms.txt for the full list of Wireflow docs.
