Predict the silence as a distribution, then decide when to speak.
A 3,137-parameter model that answers one question at 50 Hz: given
everything heard so far, how much longer will this silence last? A dynamic
program turns that distribution into a start/wait decision under an explicit,
tunable asymmetry between talking over someone and answering late.
It decides when to talk. It does not decide what to say, and it contains
no speech recogniser, no language model and no synthesiser. It is a middle layer
you drop into an existing pipeline.
Three lines
python
1from rubato import load_pretrained, CostWeights
23taker = load_pretrained(hf_repo="NagaYu/rubato-timing", weights=CostWeights.from_seconds_per_collision(2.0))4should_speak = taker.push_audio(chunk).should_speak # one 20 ms frame of float audio
seconds_per_collision is the whole configuration surface: how many seconds of
extra latency is one talk-over worth to you? Small values give an eager agent,
large values a patient one, and the sweep between them is the Pareto front below.
Framework adapters: rubato.integrations.pipecat_processor.RubatoTurnGate and
rubato.integrations.livekit_plugin.RubatoTurnDetector.
Talk-over rate against response latency: Rubato's frontier lies inside the fixed-threshold frontier
Both axes are minimised. A fixed threshold can only ever trace the outer curve -- one threshold buys one point. The inset is the band real systems ship in.
Held out on 5988 silences from speakers never seen in training (maptask, CC-BY-4.0).
covers the fixed-threshold frontier
86% of its points
covers the semantic-completeness frontier
100% of its points
latency saved at matched talk-over
84 ms vs fixed, 96 ms vs semantic
talk-over removed at matched latency
2.8 pp vs fixed, 2.1 pp vs semantic
hazard calibration (ECE)
0.0007
CRPS skill over a covariate-free hazard
+0.173
layer cost per 20 ms frame
0.26 ms median, 0.49 ms p99
At the operating point matched to a 1000 ms threshold: latency 1000 → 772 ms (95 % CI 707–831), talk-over 6.3% → 6.5% (95 % CI 5.6%–7.6%). Intervals are a conversation-level cluster bootstrap.
How it works
Discrete-time hazard.h_k = P(partner resumes in frame k+1 | silent through k, evidence up to k), from a one-hidden-layer network over a causal
feature vector: a radial-basis expansion of log elapsed silence, turn-so-far
duration and pause count, transcript completeness (gated by ASR lag, because a
real recogniser has not delivered the last word yet), terminal prosody where
audio exists, acoustic precursor cues such as in-breaths, and a running
per-partner posterior.
Optimal stopping. A backward dynamic program over the forecast survival
curve minimises alpha * P(collision) + beta * latency. Being pre-empted --
the human carries on while the agent is still silent -- costs nothing, which
is what produces the human-like behaviour: when a resumption looks likely,
waiting is nearly free; when the floor is clearly open, the agent can start at
zero gap.
Online entrainment. Conjugate posteriors over each partner's pause and gap
distributions, and an EMA of their speech rate, shrinking to a population
prior so the first thirty seconds are never worse than not adapting.
Inputs, outputs, and what it needs from you
you provide
per frame
20 ms of mono audio (any rate; 16 kHz assumed)
required
the ASR's partial transcript
optional, improves accuracy
a partner id
optional, enables entrainment
you get back
should_speak
the decision
decision.planned_onset_s
when it currently intends to start
decision.p_overlap_now
collision probability if it started this instant
prediction.future_hazards
the full predicted silence distribution
Limitations
English lexical features. The temporal and acoustic half transfers; the
completeness features were fitted on English and will mislead elsewhere.
Retrain on target-language data.
Task-oriented training data. Trained on the HCRC Map Task Corpus: two people collaborating on a route-following task. Open-domain conversation has
longer, more variable gaps; expect to retune seconds_per_collision.
Two-party only. Multi-party floor management is a different problem.
The offline protocol makes an assumption. On a floor transfer we assume the
floor would have stayed open had the agent not spoken. The alternative is
worse (informative censoring); both are reported in the repository README.
It is a timing layer, not a turn-taking oracle. It cannot know that you
paused because you were about to say something difficult.
Responsible use
Better timing makes an assistant less irritating. It also makes a synthetic voice
harder to distinguish from a person, and that is a use this model is not for.
Disclose that the system is an AI. This model is released on the assumption
that the deployment does so, in the modality the user is actually attending to
-- spoken, not buried in a settings page.
Do not use it to impersonate anyone, to make an automated caller pass as
human, or in any deployment whose value depends on the person not knowing.
Several jurisdictions require disclosure for automated calls; treat that as a
floor, not a target.
Do not use it to time interruptions strategically -- to cut a person off at
a moment calculated to stop them finishing a thought. The asymmetric cost makes
that easy to configure, which is exactly why it is named here.
The entrainment state is a behavioural profile of an individual. Keep it in
session scope unless you have a reason and consent to persist it.
Training data
maptask (CC-BY-4.0). Anderson et al. (1991), The HCRC Map Task Corpus. Language and Speech 34(4). Annotations (c) 2007 HCRC, Univ. of Edinburgh & Univ. of Glasgow. CC BY 4.0. https://groups.inf.ed.ac.uk/maptask/
No audio was redistributed in building this model.
Citation
bibtex
1@software{rubato,
2 title = {Rubato: predicting silence distributions for spoken-dialogue turn-taking},
3 year = {2026},
4 url = {https://huggingface.co/NagaYu/rubato-timing}
5}