This is an English convenience translation, provided for your information. The legally binding version is the German original; in the event of discrepancies, the German version prevails.
AI Transparency
Last updated: 24 August 2026 · Information pursuant to Art. 50 EU AI Act and Art. 13 GDPR
Coolblack is an AI-powered phone agent. We disclose which models are used for which task, what data they work with, what the models are good at — and what they are not. Callers are informed at the beginning of every conversation that they are speaking with an AI.
1. Which models are currently in use
The following list describes the app's default configuration. You can select different models in the settings at any time or store your own API keys (BYOK).
1.1 Speech recognition (speech-to-text)
- Model: Apple Speech framework + Apple SpeechAnalyzer (detection of speech pauses)
- Task: convert audio to text, detect speech pauses
- Provider: Apple Inc.
- Place of processing: on-device or via Apple servers, depending on device model and operating-system version
1.2 AI response (large language model)
Cloud mode (default for activated agents):
- Model: Google Gemini (2.5 family, EU region)
- Task: understand the caller's request, formulate a contextually appropriate response, invoke tools (calendar, contacts, reminders, call transfer)
- Provider: Google Ireland Limited / Google Cloud
- Place of processing: EU (region europe-west1, Belgium)
- Retention: at most 30 days per Google's standard; no training with your data
Offline mode (locally on the device, via Apple MLX):
- local language models (e.g. Qwen family), 4-bit quantised
- Task: as above, but without any cloud connection
- Place of processing: directly on the device
BYOK — your own API key (optional, disabled by default): This option remains inactive unless you enable it yourself. Only if you deliberately store your own API key for a third-party provider in the settings does the app use the service you selected instead of the default models — for example OpenAI (GPT family) or Anthropic (Claude). Requests then go directly from your device to that provider, on the basis of your own contract with them; their privacy policy applies. Without a key stored by you, no such transfer takes place.
1.3 Speech synthesis (text-to-speech)
Speech output depends on the voice tier you select per agent — the tier also determines where the response text is sent for synthesis. In every case: only the response text generated by the agent is transmitted, never the caller's audio.
- Standard tier (Google voices): model Google Chirp 3 HD, provider Google Ireland Limited, processing exclusively in the EU via the EU endpoint; retention at most 30 days, no training; no third-country transfer.
- Premium tier (Live voices): model Google Gemini Live — a native speech-to-speech model: the AI hears the caller directly and answers with its voice, without separate speech-recognition and synthesis steps. Provider Google Ireland Limited, processing exclusively in the EU (region europe-west4, Netherlands); the conversation audio is processed in real time, not stored and not used for training.
- Former premium voices (Cartesia — being phased out in favour of the Live voices): model Cartesia Sonic (response text only, never caller audio), provider Cartesia AI Inc. (USA), processing on the EU cluster (Frankfurt region), contractually guaranteed including failover; zero data retention, no training; safeguarded by EU standard contractual clauses, the EU-US Data Privacy Framework and a data processing agreement.
- Local tier (Apple/on-device voices): processing takes place entirely on the device — the response text never leaves the device for speech synthesis.
1.4 Tool routing
- Model: Apple Foundation Models (on-device) for entity extraction plus a proprietary rule-based router with local language-model support
- Task: decides whether the caller wants an appointment, wants a contact to be created or wishes to be transferred
- Place of processing: directly on the device
2. What data is processed
- The caller's audio (short excerpts for speech recognition)
- The transcribed text of the conversation (for the AI response and, where applicable, for creating appointments/contacts)
- Content from the client's knowledge base (FAQ, prices, opening hours) — as context for the response
- The caller's phone number and the time of the call (for CallKit and statistics)
What is stored where is explained in detail in our privacy policy.
3. Notice to callers
The agent introduces itself as an AI at the start of the conversation and points out that the call is transcribed (GDPR Art. 13, EU AI Act Art. 50). The exact greeting is configured by the client; a default wording is suggested during initial setup (e.g. "Hello, this is the digital assistant of [company]. To handle your request as well as possible, this conversation is transcribed automatically.").
3a. Task agent (outgoing calls)
Since version 3.0, users can give their agent a task by voice (e.g. a reservation); the agent then makes the call itself and carries it out. The following applies:
- Disclosure within the first second: the agent states its name, that it is a virtual voice, on whose behalf it is calling and what the call is about — before any substantive sentence.
- Own matters only: the agent only calls where the user has a genuine matter of their own. Advertising calls are excluded (Sec. 7 German Unfair Competition Act); the app requires the user's express confirmation for this.
- No recording: no audio is recorded; only a text log is created for the client.
- Bound to the facts: the agent may only state information that is part of the task or that the other party itself has mentioned. If it invents a time, number of people or price, an automatic emergency brake cuts off the utterance and forces a correction.
4. What the AI can do — and what it cannot
Strengths:
- Reliably answer standard enquiries (opening hours, prices, directions)
- Enter appointments in the calendar, create contacts and reminders
- Transfer callers to a member of staff when needed
- Multilingual — depending on the selected voice and speech recognition
Limitations / risks:
- Hallucinations: language models can invent facts that are not in the knowledge base. Verify important information (prices, appointments, legally binding statements) through measures of your own.
- Recognition errors: with dialectal pronunciation, noisy surroundings or poor call quality, speech recognition can confuse words.
- Latency: typically 0.5–2 seconds until the first response; longer in offline mode on older devices.
- No medical, legal or financial advice: the AI must not give binding information in these fields — the client configures appropriate transfer triggers for them.
5. Right to human intervention
Every caller can ask at any time to be connected to a member of staff — the app has a transfer function that the client can configure. If the caller objects to AI processing, the agent ends the conversation and offers a callback from a member of staff.
6. Responsible
Coolblack GmbH · Südfeldwiese 14 · 32107 Bad Salzuflen · Germany · info@coolblack.gmbh · Legal notice