Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · October 10, 2026 · repo

Talorysagent-infrastructureagent-securityagent-memoryprivacy

Talorys Puts a Personal AI Agent in Your Cloudflare Account. Copy Its Guardrails and Budget Its Neurons

A self-hosted assistant on Cloudflare's free tier, what its security design gets right, and the daily Neuron allowance that decides how much it can actually do.

Most personal AI agents ask you to trust a company with your notes, your tasks and everything you have ever told it about yourself. Talorys asks for something different: a Cloudflare account and one npx command. After that, the assistant runs in your account, with "no server, database or account operated by the Talorys developers," in the words of its README.

That pitch got Talorys about 185 points on Hacker News on October 10, and the repo sat at roughly 233 stars by the afternoon, with create-talorys 0.1.1 and 0.1.2 published the same day. It is MIT licensed and very young.

Two things make it worth your time. The security design makes choices worth copying, and one number on Cloudflare's pricing page sets the ceiling on everything the agent does. The README, to its credit, says so out loud.

What Talorys actually is

Talorys is a single-user assistant with streaming chat, memory, tasks, notes, projects, and one-time or recurring reminders. You install it with npx create-talorys@latest. The installer checks for Node 20.18+, confirms your Cloudflare login (or opens the authorization page), asks for an agent name and an owner password, then deploys.

The architecture is worth studying even if you never run it:

  • Your browser only talks to a Cloudflare Pages site on pages.dev.
  • Requests to /api/* go through a Pages Function, which forwards them over a service binding to a private Worker. That Worker has "no public URL."
  • The Worker hands off to TalorysAgent, a Durable Object built on the Cloudflare Agents SDK. It holds your data in SQLite, calls Workers AI, and runs alarms for reminders.
  • The model is Workers AI's @cf/zai-org/glm-4.7-flash, with streaming and tool calling.

The README is specific about what it does not create: no R2, D1, KV, Vectorize, AI Search, Workflows "or any paid service."

The guardrails worth stealing

Read docs/security.md and you find decisions that bigger agent products often skip.

The agent has no public door. The threat model "assumes the internet can reach your pages.dev URL," and the response is to make that the only door. The agent Worker deploys with workers_dev: false, and authorization is enforced in the Worker for every route except health, login and session checks.

Sessions are stored as hashes. "The database stores only HMAC(SESSION_SECRET, token); a storage leak cannot be replayed." Sessions expire after 30 days, or 14 days idle, and at most 20 are kept.

Brute force is capped twice. Five failed logins from one IP within 15 minutes lock that client, and 25 failures overall lock login globally for the window.

Deletion is gated in code, not in the prompt. The tools that remove data (delete_task, delete_note, forget_memory, cancel_automation) require confirmed: true and explicit owner confirmation. The security doc says destructive-action checks are "enforced in code, not just in the prompt."

Unattended runs get fewer powers. Scheduled AI runs receive read-only tools plus notify_owner. There is no shell, no code execution, no outbound HTTP, no credentials and no deployment tools at all.

Retrieved text is labeled as data. "The system prompt marks notes, memories, tool results and history as data."

That last one is the weakest of the six, because a label in a prompt is a request, not a wall. The other five are walls. The pattern to copy is simple: put the irreversible actions behind code, and give the version of your agent that runs while you sleep the smallest toolset you can.

The number that decides how much it can think

Here is the part people will skip and then be surprised by.

Cloudflare's Workers AI pricing page gives both the Free and Paid Workers plans "10,000 Neurons per day at no charge," resetting at 00:00 UTC. For glm-4.7-flash, it lists 5,500 Neurons per million input tokens and 36,400 Neurons per million output tokens.

Do the math on a realistic turn. Say a chat turn sends 4,000 input tokens (system prompt, recent history, relevant memories, your message) and gets 500 tokens back.

  • Input: 4,000 × 5,500 / 1,000,000 = 22 Neurons
  • Output: 500 × 36,400 / 1,000,000 = about 18 Neurons
  • Total: about 40 Neurons per turn

That is roughly 250 turns a day on the free allowance. Plenty for a personal assistant, if turns stay small. Now let a conversation grow, or let the agent chain several tool calls in one request, and the input side climbs fast. At 20,000 input tokens per call you are near 110 Neurons for input alone, and a single request with several tool steps multiplies that.

The README handles this honestly. "It is not 'unlimited free', though." When the AI allocation runs out, "chat shows a clear message and resumes after the daily reset," while reminders and digests keep working because they do not use AI. Settings → AI exposes caps on output tokens, context tokens (older history gets summarized), tool calls and reasoning steps per request, AI requests per day and scheduled AI runs per day. The Usage panel is a local estimate, and the exact count lives on the Cloudflare dashboard.

On a paid Workers plan, usage past the free allocation costs $0.011 per 1,000 Neurons, so overflow stays cheap. On the free plan, it stops.

What "your own account" does and does not mean

Talorys is self-hosted in the sense that matters for vendor lock-in: no Talorys company holds your data, and Settings → Privacy → Download backup gives you a talorys-backup.json with conversations, memories, tasks, notes, projects, settings and automations.

It is not self-hosted in the sense of nobody else touching your data. The README says so directly: "Cloudflare processes your data to provide its services," including running Workers AI inference on your chat messages "and the relevant memories included in each prompt." The processor changed from a startup to Cloudflare. It did not disappear. For many people that is an upgrade. It is still a choice you should make on purpose.

Put this into practice

If you want to try it: run npx create-talorys@latest in an empty folder, pick a long owner password, and open Settings → AI before your first long conversation. Set a daily AI request cap below what 10,000 Neurons supports so you find the limit on your terms.

If you build agents of your own, copy these four patterns this week:

  1. Give your agent backend no public URL and route everything through one authenticated front door.
  2. Make every destructive tool require a structured confirmation flag that your code checks, not a phrase the model is asked to respect.
  3. Give scheduled or background runs a separate, read-only toolset with one way to notify you.
  4. Store session tokens as HMACs so a database leak cannot be replayed.

Then do the token math. Take your model's per-token prices, measure a typical turn's input and output, and write the daily turn budget into your README before users discover it.

Honest limitations

Talorys is days old. Its first GitHub release, create-talorys@0.1.1, was tagged on October 10, and npm already had 0.1.2 a couple of hours later. Expect breaking changes.

The README says "no accounts," which means no Talorys account. You do need a Cloudflare account, and the installer will open Cloudflare's authorization page. A "Demo mode" in Settings → AI is offered as the fallback when the AI allocation runs out, with little detail on what it does.

It is single-user by design. There is no sharing, no team mode, and the security model depends on that.

The security doc does not describe encryption at rest, data deletion from Cloudflare, or a retention policy. The README does not enumerate the agent's full tool list. I did not run an independent security review, and the security claims above are the project's own.

glm-4.7-flash is a small, fast model. Expect it to be fine for tasks and notes and weaker on long reasoning. If Cloudflare changes the Neuron allocation or prices, your daily budget changes with it. The README says as much: Cloudflare sets the quotas "and can change them."

The choice in front of you

Talorys will not be everyone's assistant. What it offers is a readable blueprint: a private Durable Object, one front door, deletions behind code and a background mode with almost no power. That design costs nothing to copy.

Whether you run it or not, borrow the guardrails, and work out your Neuron budget before your agent tells you it is out of thinking for the day.

Sources: Talorys README, Talorys security doc, Cloudflare Workers AI pricing, Hacker News listing.


Medium metadata

  • Title: Talorys Puts a Personal AI Agent in Your Cloudflare Account. Copy Its Guardrails and Budget Its Neurons
  • Subtitle: A self-hosted assistant on Cloudflare's free tier, what its security design gets right, and the daily Neuron allowance that decides how much it can actually do.
  • Tags: AI Agents, Cloudflare, Self Hosted, Artificial Intelligence, Security
  • Canonical: fervorai.dev