Careers
Help build the future of human-like AI agents
Our mission is to reprogram each of the world’s largest companies using AI, potentially reaching every human in the world.
See open rolesJoin the team
Select team
Filter by location
What we value and how we act
01
Deep care for your craft
We believe great products are born from great detail. Whether you’re an engineer, designer, or operator, we expect you to set a high bar and never compromise on quality. If your work is meaningful to you, it will be meaningful at Giga.
02
Take ownership and act
Don’t wait for permission. Do the research, ask AI, and figure it out. When you see a problem, you solve it. Every person at Giga has the agency to drive change and the responsibility to deliver.
03
Operate with speed
Our advantage is speed, moving fast without breaking trust. We build and ship at a pace that surprises customers and competitors alike, while still holding ourselves to enterprise-grade standards.
04
Be customer-obsessed
We exist to give time back to humanity, starting with our customers. That means listening intently, responding quickly, and building solutions that directly move the needle for them.
05
Push boundaries
Voice has never been perfected in tech history. We’re making it happen. Curiosity and bold ideas are how we get there. At Giga, you’re encouraged to experiment, question, and push past what seems possible.
Benefits and perks
At Giga, we believe people do their best work when they feel their best.
01
Full medical, dental, and vision coverage
02
Monthly wellness benefit
03
Catered lunches, snacks, and DoorDash credits
04
Monthly commuting stipend
Films
What we’ve shipped.
- A natural voice at every turn1:37
- Introducing Giga CLI1:09
- Introducing Scout2:37
- Hallucination correction1:29
- The browser agent1:17
- Our $61M Series A2:15
- Varun and Esha with YC’s Harj Taggar35:56
- The hardest problems Giga’s AI solves10:32
- Multi-party coordination at DoorDash2:13
- The next trillion-dollar AI platform8:08
Research
giga.ai/news/stop-thinking-out-loud 4 Sep 2026
Teaching a language model to stop thinking out loud
Giga Research
Giga AI, Inc.
Abstract
In non-reasoning mode, GLM-5.2 was leaking its reasoning into the AI assistant’s response in about one conversation in a hundred. Prompt changes fixed the specific cases we had seen, but new cases kept appearing. We fine-tuned a small LoRA adapter with DPO, with a likelihood term added so that the behaviour we were not targeting did not regress. When we replayed 44,986 assistant responses from real conversations through the fine-tuned model, it produced zero leaks, against 111 for the base model.
1 The problem: reasoning traces in the assistant response
Our AI assistant uses a language model to generate the assistant’s responses and to make tool calls. For latency reasons we run the model in non-reasoning mode, so it answers directly instead of writing out a reasoning trace first.
After we moved some of our deployments onto GLM-5.2, we started seeing new issues in the conversations between the AI assistant and the user. GLM-5.2 had started adding its reasoning traces to the actual assistant response. The model would write out what the user was asking, which rule applied and what it was about to do, and only then give the reply. Here are a few representative samples of such leaks.
These issues showed up in about 1% of conversations. GPT-5.2, which we had used earlier for the same assistants, had them in about 0.03% of conversations. The leaks were also easy to reproduce. When we replayed one logged turn through GLM-5.2 forty times, it leaked on 26 of them. So this was not a rare event that happened by chance. For those turns it was the model’s default behaviour.
Table 1: Replays of real conversations: 44,986 assistant turns over 4,359 conversations, from five deployments, one reply per turn from each model.
| Model | Turns | Leaks |
|---|---|---|
| GLM-5.2, base (replayed) | 44,986 | 111 |
| GLM-5.2 + DPO (LoRA) | 44,986 | 0 |
2 What prompting could and could not fix
We started with prompt engineering. Every leak we observed happened at a point where the model had to make a decision while writing its reply: whether to offer a transfer, whether the other side was a voicemail system, whether to switch language. At each of these points we added an instruction to the prompt that told the model exactly what its reply should look like. For example, at the point where the assistant decides whether to offer a transfer, the prompt said that the entire reply should be one short confirmation question, and gave an example of one. This fixed the leak at that point almost completely. We also added a short section to the prompt that stated the rule that everything the model writes is shown to the user, with a pair of WRONG and RIGHT example replies, and we repeated that rule in one sentence at the very end of the prompt as a reminder.
Prompting was patching each issue we encountered, and it did not provide a general fix for the problem. A leak that fired on 165 of 165 replays in one conversation state went to zero once we scripted that state, and then a neighbouring state started leaking. After every fix there was still a floor of roughly half a percent to one percent of conversations, and the remaining leaks were always at a decision point we had not scripted yet, or on users who explicitly asked the model to explain its reasoning, or on language-switch decisions.
3 Finding the root cause
To understand what was going on, we turned reasoning mode on and replayed the same leak-prone contexts. With reasoning on, the leaks disappeared (0 of 36 replays, compared with 15 of 29 with reasoning off), the tool syntax leaks disappeared too, and tool calls became more reliable. The cost was latency: time to first token went up by 1 to 4 seconds on decision turns. Our assistants have to answer in real time, so we could not simply turn reasoning on for every turn.
This experiment also explained the defect. The model has learned to reason before it answers. The non-reasoning chat template asks it to skip that step by pre-filling an empty think block at the start of the response.
1
Sep 20260 leaks in 44,986 replayed responses, against 111
giga.ai/news/synthetic-voice-detection 13 Aug 2026
Real-Time Synthetic Voice Detection Comes to Giga
Giga Research
Giga AI, Inc.
Abstract
Today we are launching real-time synthetic voice detection on the Giga platform. Every call handled by a Giga voice agent can now be screened, as the conversation happens, for signs that the speaker on the other end is not a person but a machine.
1 Why this matters now
Voice used to be the strongest trust signal a business had. If someone called in and sounded like a customer, they almost certainly were one. That assumption has quietly stopped being true, because modern text-to-speech systems can clone a convincing voice from just a few seconds of reference audio [1], and they are cheap, fast, and available to anyone.
The consequences are no longer hypothetical. Criminals used a cloned executive voice to authorize a fraudulent transfer of $243,000 as far back as 2019 [2], and in 2024 a finance employee in Hong Kong wired roughly $25 million after a video call staffed entirely by deepfaked colleagues [3]. Deloitte’s Center for Financial Services projects that generative AI could push fraud losses in the United States to $40 billion by 2027 [4].
Figure 1: Three amounts on one money ruler; every major step is ten times the last. The 2019 and 2024 amounts are single documented payouts. The 2027 amount is a projected national annual total, a different kind of number, marked as such.
Contact centers face a newer, quieter version of the same attack. Instead of one dramatic heist, bad actors can point AI voice agents at customer support lines and run them as persistent callers: placing many routine calls, probing the refund policy, accepting escalating politely, and retrying across interactions until a rare policy edge case pays out. Any single call looks routine, but taken together this is industrialized fraud running at a scale no human caller could sustain. We built this product because our customers, including some of the largest consumer brands in the world, asked us to help them stop it.
2 Why synthetic voice detection is hard
The obvious approach sounds simple enough: train a classifier that separates real speech from generated speech. The research community has been working on exactly this problem for a decade through the ASVspoof challenge series [5], and its results explain why the problem resists easy solutions.
The first obstacle is generalization. Detectors that score near-perfectly against the synthesizers they were trained on degrade sharply when they meet voices from new, unseen models. Müller et al. found that error rates of published detectors increased severalfold when evaluated on in-the-wild audio instead of curated benchmarks [6]. New voice models ship every month, so a detector that memorizes the artifacts of today’s synthesizers is obsolete on arrival. It has to learn what makes speech human rather than what makes one particular model sound fake.
The phone network itself works against detection. Much of the research literature is built on clean, high-sample-rate audio, while real calls are narrowband, compressed by lossy codecs, and degraded by packet loss and background noise. The ASVspoof 2021 evaluation showed that these transmission effects can substantially weaken common detection systems [7], since many of the spectral artifacts that give away synthetic speech in the lab live in exactly the frequencies the phone network throws away.
Timing adds a further constraint. A forensic model that analyzes a complete recording after the fact is useful for audits, but it cannot stop a live fraud attempt. In a contact center, detection has to run on streaming audio and produce a confident signal within the first moments of a conversation, while the caller is still mid-sentence, without adding any latency to the call.
1
Aug 2026Every call screened for a synthetic voice as it happens
giga.ai/hallucinations 7 May 2026
Real-Time Hallucination Correction at Zero Latency Cost
Esha DinneRishi AlluriArnab Maiti
Giga AI, Inc.
Abstract
We reduced voice agent hallucination rates from 4-5% to less than 1% in production without adding latency. LLMs generate text faster than humans can speak it. That speed gap is where we run detection.
1 Why Hallucinations in Voice Are More Dangerous
A caller asks a voice agent about their copay. The agent says, “$0.” The real amount is $40. The caller books the appointment, drives across town, and learns the truth at the front desk.
Voice makes this worse than text for two reasons.
You cannot use reasoning models for generation. Voice requires roughly one-second time to first byte, or the conversation feels broken. Reasoning models think before responding and hallucinate less, but their first token takes several seconds, which is too slow for voice. This forces voice systems to use non-reasoning LLMs, which are quicker but hallucinate more. The constraint that makes voice feel natural is the same constraint that makes it less accurate.
Spoken errors bypass verification. In text, a wrong answer stays on screen. The user can re-read it, question it, or check another source. In voice, the same answer arrives as a confident statement and then disappears. Listeners use a speaker’s tone, rhythm, and emphasis to judge how certain and honest they seem, and those vocal cues can influence both perceived reliability and verbal working memory (Goupil et al., Nature Communications, 2021). The caller is more likely to act before verifying [1].
Figure 1: Production results, measured on live traffic across 1.2 million conversational turns, 30-day rolling. Hallucination rate fell more than 70%, from a 4–5% baseline to under 1% absolute, with false positives below 0.3%.
These compound. The model hallucinates more because it cannot reason. The caller trusts it more because it sounds confident.
2 Why You Can’t Just Verify Before Speaking
The standard fix: verify the response with a second model before the caller hears it. In text, this works. In voice, it adds 3 to 4 seconds of latency to every turn. TTFB goes to 3+ seconds. The caller hangs up.
You pay this on every turn, even on the 96% of turns where the model didn’t hallucinate.
3 The Core Insight: Throughput > Speaking Speed
LLMs generate text far faster than voice models can speak it.
The right comparison is not raw tokens per second alone. A voice system first pays time to first token, then streams the remaining output. Using rough planning numbers, a recent low-latency model might take ~600ms to emit the first token and then stream at ~75 tokens per second. A 30-word response is roughly 40 output tokens, so the full text is available in about a second. Speaking those same 30 words takes about 10 to 12 seconds. That several-second gap between what has been generated and what has been heard is where we run detection.
This is not a small window. Even after including the initial latency, the full text usually exists many seconds before the caller finishes hearing it.
1
May 2026<1% of turns hallucinate, from 4–5%, with no added latency
giga.ai/news/building-a-prefill-first-classifier-llm 4 Dec 2025
Building a Prefill-First Classifier LLM
Giga Research
Giga AI, Inc.
Abstract
Most classification problems do not require an LLM. However, for problems with long, messy, or conversational inputs, a classifier LLM is the best approach for capturing semantic meaning, ensuring accurate decision-making. Because we are building a voice AI for customer support, we needed a classifier LLM for multi-turn interaction logs, system messages, and policy texts.
1 TL;DR
Classifier LLMs have a unique topography. Specifically, they have a very asymmetrical input-to-output token ratio. Inputs often extend up to 10,000 tokens (sometimes, even 60,000+ tokens); conversely, outputs rarely cross a dozen tokens. A fair share of LLM research is focused on the decode stage, when the LLM generates output tokens. However, for classifiers, the focus is instead on prefill, when the LLM reads tokens.
Last quarter, we decided to build a classifier LLM atop an open-weight instruction model that satisfied a few internal design requirements. We improved our system’s latency, precision, and recall compared to our build that used an off-the-shelf model. This blog discusses how we accomplished all of this with a sub-12 hour fine-tuning time.
Figure 1: Latency at P50, P90 and P99. Our fine-tuned Qwen had a P50 latency of 150ms, over 3x faster than GPT-4.1 mini’s latency of 450ms.
2 Design Goals
At Giga, we build voice AI agents for customer support. Because our agent directly interacts with our customers’ customers, we needed a product that could work with diverse clusters of end users. This translates to a few non-negotiable technical constraints:
Low latency: Poor latency can make verbal conversations feel unnatural; humans are used to quick responses and get frustrated if they perceive that they’re talking to a robot. Our classifier needed to accommodate tight time budgets. Additionally, because inbound customer support lines surge during an outage, latency cannot collapse during bursty loads.
Modularity: GigaML’s customers have different customer service needs, so we needed to fine-tune behavior without retraining the base model for each customer. Otherwise, we’d be faced with a computationally expensive runtime for every customer and whenever a customer’s needs evolve. Instead, we required a tenant-host system that would allow for hot-swapping different trained variants.
Cost control: We valued having a predictable per-request cost. Because classification is a prefill-heavy problem, our backend had to be optimized for prefill efficiency, not decode efficiency.
3 Our Approach
For our base model, we chose Qwen3-8B, a modern ~8B-parameter open-weight instruction model. There were two primary reasons why Qwen3-8B was the ideal choice:
Capacity and latency balance: 8B parameters are enough to handle long contexts and subtle label boundaries. It’s also small enough to hit aggressive real-time benchmarks by optimizing prefill.
Ecosystem: Qwen3-8B has a great tokenizer and instruction obedience out of the box. It also supports fine-tuning, quantization, and modern inference stacks.
We chose to keep Qwen3-8B as a frozen base; we do not fine-tune the model directly, keeping all of the original weights. Instead, we use lightweight LoRA (low-rank adaptation) adapters to create an efficient low-dimensional representation of fine-tuning updates.
There are many benefits of using LoRA adapters that closely align with our design criteria:
Isolated by design: Each adapter has its own set of weights, each tuned to a single customer success policy. By changing (or breaking) one adapter, we don’t affect the behavior of another. We could have hundreds of available adapters at scale.
1
Dec 2025150ms at P50, over 3× faster than GPT-4.1 mini
giga.ai/news/fluent-by-design 5 Oct 2025
Fluent by design
Giga Research
Giga AI, Inc.
Abstract
In customer conversations, clarity isn’t just about what’s said. It’s about being understood. As Giga agents expand to serve our global user base, we’ve been determined to always answer the following question: Can AI transcend language? That’s the idea behind agent multilinguality, every caller should be able to interact comfortably in their preferred language.
1 The challenge
Most customer-support systems rely on manual routing or translation layers to handle different languages. These approaches slow down response times and introduce friction, especially in real-time voice experiences.
For marketplaces like DoorDash, where every second matters, language barriers can break the flow of a delivery, an order, or an escalation. A great AI agent shouldn’t need to pause to translate. It should just understand.
2 What we built
Multilinguality allows agents to detect, remember, and respond in the customer’s preferred language across voice, chat, and SMS.
When a first-time caller connects, the system identifies their language and offers a quick choice. From that moment forward, every interaction happens in that language, with no prompts or transfers required.
Agents can now speak and reason across 90+ languages.
3 How it works
Table 1: How it works.
| Component | What it does |
|---|---|
| Language Identification | Real-time detection of the caller’s spoken language. |
| Preference Memory | The system stores each user’s language choice for future interactions. |
| Dynamic Routing | Calls automatically reach agents trained or optimized for that language. |
| Multilingual Reasoning | Giga agents can switch languages mid-conversation whenever necessary (for example, English ↔ Spanish). |
| Analytics and Insights | Language data is tracked across tickets, enabling multilingual reporting and performance analysis. |
4 Why it matters
For global operations, language has always been the silent bottleneck in automation.
By removing it, Multilinguality unlocks four key advantages:
Deeper Understanding: Agents grasp nuance, intent, emotion, and context across languages, reducing confusion and friction.
Higher Resolution: Fewer transfers. Faster responses. Every call starts with comprehension, not translation.
Happier Callers: People feel understood from the first word, building trust, connection, and satisfaction across cultures.
Cultural Fluency: Our voice models adapt to regional accents and tone, making every interaction feel local and human.
Multilinguality is part of a broader vision for adaptive voice intelligence, systems that listen as well as they respond.
As Giga agents continue to learn across new markets and domains, we’re building toward a world where customer experience is defined not by where you are or what language you speak, but by how quickly you’re understood.
1
Oct 202590+ languages, with no transfers
giga.ai/news/scaling-rag-to-100k-documents-without-lag 25 Aug 2025
Scaling RAG to 100K+ Documents Without Lag
Giga Research
Giga AI, Inc.
Abstract
The challenge comes when the corpus grows into the hundreds of thousands or millions of documents. At that scale, retrieval latency becomes the difference between a seamless, human-like interaction and an experience that feels sluggish or broken. This article breaks down how to architect, optimize, and operate a RAG pipeline that scales to massive corpora without lag, using real-world benchmarks, architectural patterns, and techniques applied by enterprise teams in production.
1 The Scaling Challenge
Retrieval-Augmented Generation (RAG) is the bridge between large language models and the fresh, domain-specific data they need to deliver accurate answers. It’s the reason an AI can answer a question about your latest compliance policy or sales report, even if those documents were added minutes ago.
Scaling RAG is not just about adding more data. It’s about keeping latency low across the entire pipeline:
Embedding search latency: At scale, brute-force search is impractical. Approximate Nearest Neighbor (ANN) algorithms like HNSW and IVF+PQ make large-scale vector search feasible, but require tuning to balance recall, memory footprint, and speed.
Context assembly overhead: Fetching, re-ranking, and merging retrieved chunks into a prompt can be a hidden bottleneck, especially with long contexts or complex ranking logic.
LLM processing time: Even perfect retrieval loses its impact if the Time-to-First-Token (TTFT) is high. Token processing costs grow with context size, so inference optimizations matter as much as retrieval speed.
The takeaway: performance tuning must be holistic. Gains in one stage can be erased if another stage lags.
2 Architecting for Speed
2.1 Choosing the Right Vector Database
Your vector database is the foundation of your RAG latency profile. Performance varies widely between systems, both in P99 latency and queries per second (QPS).
Table 1: Consolidated benchmark table: P99 latency and QPS (1M static).
| Database (Config) | P99 (ms) | QPS |
|---|---|---|
| ZillizCloud-8cu-perf | 2.5 | 9,704 |
| Milvus-16c64g-sq8 | 2.2 | 3,465 |
| OpenSearch-16c128g-force | 7.2 | 3,055 |
| QdrantCloud-16c64g | 6.4 | 1,242 |
| Pinecone-p2.x8-1node | 13.7 | 1,147 |
| OpenSearch-16c128g | 13.2 | 951 |
Key insights from the benchmark:
ZillizCloud: 2.5 ms P99 latency and ~9,700 QPS — strong for high-throughput, low-latency needs.
Milvus: 2.2 ms latency with ~3,465 QPS — competitive on latency, slightly lower on throughput.
OpenSearch (force merge): Higher latency (~7.2 ms) but still competitive QPS for certain workloads.
Pinecone, Qdrant, OpenSearch (standard): Trade-offs in raw speed for operational simplicity or ecosystem integration.
2.2 Streaming Performance Matters
Static benchmarks only tell half the story. Many real-world RAG applications run on constantly updated corpora, so the database must handle concurrent reads and writes without collapsing under ingestion pressure.
Highlights:
ZillizCloud shows strong resilience, maintaining ~1,860 QPS at 1,000 rows/s ingestion.
Qdrant degrades minimally under ingestion.
Pinecone and OpenSearch experience sharper drops as ingestion rates increase.
3 ANN Algorithm Choice
HNSW: Excellent speed-recall trade-off, versatile across datasets, but memory-heavy.
IVF+PQ: Lower memory footprint, fast for certain distributions, but more sensitive to parameter tuning.
The choice depends on embedding dimensionality, recall tolerance, and available hardware.
1
Aug 2025100K+ documents, retrieved without lag
Join our team
Ready to build the next generation of reasoning AI?
Join the team shaping how enterprises around the world communicate, solve problems, and scale support with human-like intelligence.
See open roles