Log in to AgentMeter
AgentMeter measures how efficiently open-source language models run a fixed 4-agent analysis task — latency, GPU memory, tokens, cost and energy. It is a measurement tool, not a threat-detection product.
Benchmark the resource efficiency of LLMs on your network flows
AgentMeter measures how efficiently open-source language models run a fixed 4-agent analysis task: latency, GPU memory, tokens, cost and energy — plus accuracy as context when your data has labels. It is a measurement tool, not a threat-detection product, and it makes no security decision about your network.
Home
Measure how efficiently two open-source LLMs run the same 4-agent task on your flows: latency, memory, tokens, cost and energy, on the same GPU.
Start a new benchmark
Upload a CSV or PCAP (or give a link), pick up to two models, and run. Data preparation runs in the background while you choose.
No sessions yet
A session is one benchmark run of up to two models on your flows — run your first one to see it here.
New benchmarkRecent sessions
All sessionsNew benchmark
Three steps. Your data is prepared in the background while you pick the models.
- 1Data
- 2Models and settings
- 3Run
Prepared on this server
Your data
PreparingModels
Pick up to two models. Both run the same flows on the same GPU, one at a time.
Select 1 or 2 models.
Settings
Starting
…You can close this page — the run continues on the server. Find it again under Sessions.
Sessions
Every benchmark run. Sessions are non-validated user runs and never merge into the locked study.
Session · non-validated user run
Session
Compare two sessions
Every efficiency metric side by side. A comparison is like-for-like only on the same prepared set, the same GPU and the same settings.
Leaderboard
Models ranked across sessions — but only against models measured on the same prepared set and the same GPU. Each group is its own ranking.
Settings
Theme is saved in this browser.
Backend
A new backend address can be given once with ?api=<url> (remembered in this browser; ?api=reset forgets it).
Admin
Accounts (at most five), credential status and system limits. Secrets are never shown here — only whether they are set.
Users
Credentials
Set on the GPU backend (Hugging Face token, Vast API key) or as a Worker secret (Telegram bot token). Status only.