Only 8B active parameters and a much smaller KV cache, so it triages thousands of alerts quickly and cheaply.
A multimodal MoE that activates only 8B parameters per token and shrinks the KV cache about 4x.
Input: text, image
Harnesses in Runtime
Risk and fraud
For transaction monitoring, our top pick is DeepSeek V4.1 Flash from DeepSeek. Only 8B active parameters and a much smaller KV cache, so it triages thousands of alerts quickly and cheaply. GLM-5.3 Flash is a close second, and gpt-oss-120b rounds out the list.
Monitoring agents triage alerts from rules engines and SIEMs, enrich them with account and transaction history, and escalate the real ones. Volume is high and most alerts are false positives, so speed and cost per alert matter most, with enough reasoning to explain each call.
Updated
Only 8B active parameters and a much smaller KV cache, so it triages thousands of alerts quickly and cheaply.
A multimodal MoE that activates only 8B parameters per token and shrinks the KV cache about 4x.
Input: text, image
Harnesses in Runtime
18B active parameters with strong agentic skills for the enrichment queries behind each alert.
A natively multimodal GLM that Z.ai says beats GLM-5.2 at one-tenth the price.
Input: text, image
Harnesses in Runtime
Apache 2.0 and fits on a single 80GB GPU, a simple way to run triage entirely inside your network.
OpenAI's open-weight reasoning model, small enough to run on a single 80GB GPU.
Input: text
Harnesses in Runtime
Turn the steps your team already follows into a skill the agent reads before every run.
Run two or three of these models against past cases. Evals show which one passes at the lowest cost.
The agent works on its own computer in your VPC, with role-based access and an approval before sensitive actions.
For transaction monitoring, our top pick is DeepSeek V4.1 Flash from DeepSeek. Only 8B active parameters and a much smaller KV cache, so it triages thousands of alerts quickly and cheaply. GLM-5.3 Flash is a close second, and gpt-oss-120b rounds out the list.
Low cost and latency per alert. Good tool use for enrichment queries. Easy to run at volume in your own cloud. Test two or three candidates against your own past cases before you commit, and keep a human approval on any step that moves money or changes a customer's status.
Yes. Download the weights and serve them with vLLM or SGLang in your own cloud, or use Amazon Bedrock or Google Vertex AI in your account. With Runtime, the agent's computer also runs in your VPC, so card data and PII stay in your environment.
DeepSeek V4.1 Flash (MIT), GLM-5.3 Flash (MIT), gpt-oss-120b (Apache 2.0) are permissive and fine for commercial use.