Clapilot-Agent Heartbeat
Recurring proactive checks for users, teamchat, and opt-in self-acting specialized agents.
The heartbeat lets Clapilot check in proactively instead of only answering when asked. User and teamchat scopes may post at most one proactive chat message — by default none. A third agent scope powers opt-in Self Acting Agents: each enabled specialized agent periodically works its own shared Aufgaben board, reads the Team Chat, and communicates through explicit Team Chat posts (claims, results, questions, follow-ups on blockers, a daily standup) — its final heartbeat reply itself is never posted. This is the v2 implementation; the earlier heartbeat subsystem was removed and its legacy native_heartbeat_* settings columns are intentionally ignored.
What a heartbeat run does
- The
clapilot-agentruntime executes an agent turn in a dedicated session (heartbeat:user:<userId>,heartbeat:teamchat, orheartbeat:agent:<specializedAgentId>) with the scope-specific instructions and context. - The agent decides whether anything genuinely needs attention right now, using tools to verify facts during the run. Findings from earlier runs must be re-verified, never repeated blindly.
- The final reply is gated by the
NO_MESSAGEtoken — a sentinel word the model is instructed to reply with when there is nothing worth posting:- Reply is exactly
NO_MESSAGE(or the token plus trivial filler under 20 characters) → nothing happens: no chat message, no notification. - Reply contains real content → it is posted to the target (personal chat via
main_sessiondelivery, or the teamchat room) withmessage_origin = 'assistant_heartbeat'. - Reply contains substantial content alongside the token → the token is stripped and the content is delivered.
- Reply is exactly
Configuration
Settings UI: Profil → Heartbeat (per user) and Admin → Teamchat-Heartbeat (admin only). Both panels offer a manual "Run now" trigger that executes immediately, reports whether a message was posted or suppressed, and does not shift the regular schedule.
| Field | Meaning |
|---|---|
enabled | Master toggle. Off means the runtime disables the job entirely. |
intervalMinutes | How often the heartbeat runs (5 min – 7 days, default 60). |
instructions | Free-text operator instructions executed on every run (max 8000 chars). |
sessionId (user scope) | Target chat session; null = main chat. History is read from and messages are posted to this session. |
roomId (teamchat scope) | Teamchat room; null = default clapilot-members (#general). |
| Teamchat identity | The teamchat heartbeat runs as the global_team_service service principal with no human user (team automations must not retain a user's runtime identity; migration 231_remove_legacy_team_automation_user_scope.sql strips the legacy actingUserId). Workspace-global tools (tasks, documents, wiki, customers) work normally. Personal-first tools such as calendar_* return CLAPILOT_CALENDAR_USER_SCOPE_REQUIRED with a localized explanation, and the teamchat prompt tells the agent not to report that as an outage. Appointment reminders belong to each user's personal heartbeat (heartbeat:user:<userId>). |
historyLimit | How many recent messages from the target chat are included as context (0–50, default 15; 0 = none). |
model | Optional model ref override for heartbeat runs (picker backed by the same model list as the chat UI); null = configured primary/fallback model routing. |
dndStart / dndEnd | Optional quiet-hours window ("HH:mm", both required; overnight windows like 22:00–07:00 supported). While inside the window, scheduled runs are skipped entirely (no agent turn, no tokens) and deferred to the window's end. The manual "Run now" trigger deliberately bypasses quiet hours. |
dndTimezone | IANA timezone the quiet-hours window is interpreted in (default Europe/Berlin; not exposed in the UI). |
Storage:
- Per user:
user_profiles.heartbeat_config_json(migration 180) - Teamchat:
app_settings.teamchat_heartbeat_config_json(migration 180) - Self-acting specialized agent:
specialized_agents.self_acting_config_json(migration 285). This scope usesintervalMinutes,instructions,teamChatScanEnabled(defaulttrue),reportChannelId(defaultnull),historyLimit(earlier home-channel messages shown as context, default 15, at most 10 when a scan cursor exists),dailyStandupEnabled(defaulttrue),actorUserId(defaultnull, admin-only), and quiet hoursdndStart/dndEnd/dndTimezone. Quiet hours default to 00:00–08:00 Europe/Berlin when the config never set them; an explicitnull(the switch turned off in the editor) lets the agent run around the clock, and an incomplete window counts as off. Scheduled runs inside the window are skipped and deferred to its end exactly like user heartbeats; the manual run trigger bypasses quiet hours. The board triage always loads all open board tasks (up to 50). PATCH requests mergeselfActingConfigkey by key over the stored config, so keys a client does not send are kept. Standup state lives inspecialized_agent_heartbeat_state(migration 338).
Self-acting specialized-agent scope
Enabling Self Acting Agent on an enabled specialized agent creates or reuses a dedicated shared Aufgaben board and declaratively reconciles one heartbeat job for it. Each run reloads the full specialist definition, uses the specialist's default model and permission envelope, and passes personalMemoryEnabled into executeRun, so its existing isolated personal memory applies automatically. Approved personal memories are also projected into that specialist's isolated Knowledge Graph (scope_key='agent:<id>') and are reachable through the knowledge_* tools only during its own runs. The prompt contains the configured heartbeat instructions, the triaged board tasks described below, and that agent's watchlist.
Task triage, blocker protocol, and follow-ups
Every run loads all open tasks of the dedicated board together with three activity signals per task: the agent's own last comment (aufgaben_kommentare rows with autor_typ='agent' and the agent's name), the last comment by anybody else, and the number of agent comments since that external comment (agent_comments_since_activity). The runtime sorts the tasks into prompt sections instead of listing them all as work to continue:
- Actionable — every task whose status category is
openorin_progress. The agent does the work with tools now, verifies it with tool output, and finishes each task touch with exactly one comment (what was done, evidence, next step). - Waiting with new activity — a
waitingtask that the agent never commented on, that someone else commented on or changed after the agent's last comment, or whose due date /wiedervorlage_atpassed since that comment. The agent acts on the reply, comments once, and tells the person in Team Chat what happens next. - Follow-up due — a
waitingtask whose follow-up interval elapsed without any response. The interval follows the backoff ladderAGENT_TASK_FOLLOW_UP_HOURS= 24h after the blocker comment, 48h after the first follow-up, then every 72h until someone responds (the step isagent_comments_since_activity). The agent first retries the blocked step once (access or tokens may have been fixed) or looks for another route; if still blocked it posts a Team Chat message addressed with@namestating what it needs, why, and what it will do next, plus one comment on the task. That comment advances the ladder. - Parked — every other
waitingtask, listed compactly with the time until its next follow-up. It returns automatically when someone responds, a due date passes, or the next follow-up is due.
The prompt carries a self-unblock rule and the blocker protocol, and states that these rules replace older heartbeat policy: notes in the agent's memory or watchlist from earlier runs that restrict it to certain report events, forbid chat replies, or say not to re-ping someone are ignored as outdated policy while their facts stay usable. A blocker is a problem to solve first: before asking a human the agent tries other tools (tool_catalog_search), wiki/documents/memory/knowledge search, re-checks whether an access problem was fixed, does the part it can do, or asks another agent in Team Chat with a single @handle mention (agent-to-agent reactions require the channel's opt-in agent_to_agent_enabled). Only what needs a human decision, credential, or input goes to a person: one task comment naming the input and the person, the task moves to the waiting status key, and one Team Chat post addresses that person. The follow-up ladder then keeps chasing it.
aufgaben_add_comment is callable from self-acting heartbeat sessions (heartbeat:agent:<agent id>) and mention sessions (agent:<handle|id>:openai-user:…) for tasks on the agent's own self_acting_board_id or on any shared board (aufgaben_boards.visibility_scope = 'shared'), provided the enabled agent's tool allowlist includes it. This lets agents ask assignees about stale tickets or document progress on team boards. Private boards stay off limits (SELF_ACTING_TASK_BOARD_MISMATCH). Comments are stored as autor_typ='agent' with the agent's name and identical text by the same agent is deduplicated. Creating, updating, and moving tasks already works on every shared board.
Team Chat awareness
When teamChatScanEnabled is true, the heartbeat reads new messages from every non-archived Team Chat room to which the specialist is invited. chat_room_specialized_agents.heartbeat_scanned_at stores the cursor per room/agent pair:
- Member messages (
nachricht) and agent/assistant replies (antwort, clipped to 600 characters) are both shown; replies are labelled(agent reply)and the agent's own postsyou (earlier post). Member messages that@mentionthe agent are flagged(mentions you). - Each room contributes up to 60 rows per heartbeat. The cursor advances past every examined row, including agent replies without member text, and is read and written with full microsecond precision (
to_json(created_at)). Earlier versions only advanced on member messages and truncated to milliseconds, which could pin a cursor on one message forever. - Cursors older than 48 hours (first scan, long pause, or the old stuck-cursor state) are clamped to the last 48 hours instead of replaying the backlog.
- Up to
historyLimitearlier messages of the home (report) channel before the unscanned window are included as read-only context so the agent sees whole threads.
Matching work is claimed by calling aufgaben_create_task on its dedicated board with source_type="team_chat" and source_id=<message id>; the native path returns the existing task when that source is already claimed, and a partial unique index on Team Chat source ids closes concurrent-agent races. Direct @mentions are still answered live by the mention path, so the heartbeat prompt tells the agent not to answer the same message twice but to follow through on the requested work.
Heartbeat runs always receive team_chat_post_message and team_chat_read_messages, so admins do not need to add them to the allowlist. team_chat_read_messages returns recent messages of public channels, private channels the acting user belongs to, and channels the agent is invited to (never direct messages), oldest first, with paging via before and an optional text query.
Communication and daily standup
With a reportChannelId (non-archived channel_public or channel_private), the prompt makes the Team Chat the agent's primary output. It posts when it claims or starts work, finishes something (result and where to find it), needs input (@name), a follow-up is due, it can answer a question, or proactive work found something worth sharing. It must not post unchanged state or bookkeeping and sees its own last five posts (matched through message_meta.assistantAgentId) to avoid repeats. Posts are brief (1–4 sentences) and written in the team's chat language. Status posts carry the specialist's identity meta (assistantAgentId / handle / name / profile image), so they render as the agent in Team Chat.
When dailyStandupEnabled is on, the first heartbeat on a weekday at or after 09:00 Europe/Berlin asks the agent for one standup post in the home channel: finished since the last standup, next up, and blockers with @name. The day is marked done in specialized_agent_heartbeat_state.last_standup_on once the run made a successful team_chat_post_message call; otherwise the next heartbeat asks again.
Standing duties from the instructions (stale-ticket checks, bookmark scans, open requests) must advance by at least one item per heartbeat; the prompt forbids self-invented slower schedules such as weekly sweeps or self-set snoozes. Without configured instructions the prompt tells the agent to derive proactive work from its role and system prompt: open questions in Team Chat it can answer, unassigned or stale tasks in its area on shared boards, and follow-ups nobody picked up.
Acting user (actorUserId)
Heartbeats run under the global team service principal and have no user by default, so user-scoped tools (browser_*, x_get_bookmarks, video_studio_* voiceover, emails_*) fail with missing-user-context errors. An admin can set Handeln als Nutzer (actorUserId) on a self-acting agent. The heartbeat then passes that user as actorUserId into executeRun. The service-principal identity is kept for memory, media ownership, and Team Chat posting, while tool calls resolve scopedUserId to the configured user: their browser profile, X connection, mailbox, and Video Studio access, and outbound-action approvals are routed to that user. The tool proxy only honours the persisted agent_session_state.actor_user_id for heartbeat:agent:<id> sessions while it still equals the enabled agent's self_acting_config_json.actorUserId, so a cleared or changed setting takes effect immediately. The value can only be set by an admin in the settings UI (create/update APIs validate that the user exists and is not a pending invitation); agent tools (specialized_agents_create / _update) reject it with SELF_ACTING_ACTOR_ADMIN_UI_ONLY. This is a deliberate, opt-in exception to the rule that team automations carry no human identity (migration 231); task visibility stays shared-only because task scoping still uses the userless identity.
Run outcome
The run is classified as an automation through clientContext.automationId = "self-acting-agent:<agentId>", so autonomous completion judging and caveat safeguards apply. The heartbeat's final reply is never delivered to chat: the runtime only applies WATCH: / RESOLVE: directives, marks due watch items checked, advances scan cursors, records the standup, and logs the run outcome ([heartbeat] self-acting agent run completed with task bucket counts, scanned messages, mentions, chat posts, standup and actor flags). Successful runs persist their tool calls in agent_runs.tool_calls (last 200, 500-character previews), so the run log shows real tool counts even when the session's event window no longer covers older runs. Its per-agent Knowledge Graph stays excluded from workspace builds, Wiki, Dreaming, user runs, and every other specialist.
Watchlist
Each heartbeat scope has a persistent watchlist — items to re-check on later runs ("waiting for a reply from X", "check whether invoice 4711 was paid"). User items are keyed by user_id, teamchat items have no user, and agent-scope items are keyed by specialized_agent_id. The model manages the list through directive lines in its final reply, which the runtime parses and strips before delivery handling (the same final-reply-protocol pattern as the automation notify prefix):
WATCH: <what to check>— keep watching, re-verify on every runWATCH[24h]: <what>/WATCH[3d]: <what>— keep watching but snooze rechecking for that durationRESOLVE: <short-id>— remove an open item (ids are shown in the run prompt)
NO_MESSAGE plus WATCH: lines therefore stores items without posting anything. Open items are injected into every run prompt with the explicit rule that they are reminders to re-verify, never facts to repeat (the v1 stale-notes failure mode). Guardrails: max 20 items per scope, 500 chars each, mandatory auto-expiry after 14 days, duplicate adds refresh the existing item instead of duplicating it, and an item that was presented as "recheck now" WATCHLIST_MAX_CHECKS (8) times without being resolved is expired automatically (check_count, migration 328; the prompt shows the count per item). A RESOLVE: and a matching WATCH: in the same reply store the new item instead of updating the deleted row. For self-acting agents, WATCH: lines that mention a task id from the agent's own board are discarded before storage: board tasks are tracked by the board (status, comments, due dates), never by a second list the model would re-verify every run. Storage: heartbeat_watch_items (migration 182). The settings panels show the current watchlist and let users delete items directly (DELETE /api/heartbeat/watchlist?scope=...&id=...).
Scheduling model
The runtime module services/clapilot-agent/src/jobs/heartbeat.mjs declaratively reconciles agent_jobs rows (job_type = 'heartbeat') from the stored configs: instead of syncing at save time, a poll loop (default every 30 s, CLAPILOT_AGENT_HEARTBEAT_POLL_MS; reconciliation on every second tick) continuously makes the job table match the stored configs. Enabling, disabling, or editing a heartbeat therefore takes effect within about a minute. Due heartbeats run concurrently, up to CLAPILOT_AGENT_HEARTBEAT_MAX_CONCURRENT_RUNS (default 4) at a time, so one long self-acting agent run no longer delays every other agent, user, and teamchat heartbeat; a per-job advisory lock and an in-process in-flight set keep each job single-run. Config edits never postpone an already-due run; shortening the interval takes effect immediately.
Heartbeat provider calls stream internally so the provider deadline measures inactivity instead of total response time. They retain the background-workload retry and fallback policy, including per-provider and per-model timeout overrides for slower self-hosted deployments. A successful side-effecting tool boundary is checkpointed before another provider turn can run, so a later stall continues from durable progress rather than replaying the completed action.
The operational heartbeat SLO is fewer than 1% failed scheduled runs in a rolling 24-hour window. A timeout regression is not considered resolved until the affected production heartbeat has completed a full 24-hour observation window without the reported timeout signature; implementation and merge alone are not closure evidence.
Runs are guarded by PostgreSQL advisory locks (application-level locks held only for the duration of the run), so concurrent pollers can never double-post. Failures are recorded on the job row (last_error, last_delivery_status = 'error'), surface only in the settings UI status panel, and the next attempt simply happens at the next interval — no retry storms, no failure messages in chat.
If the runtime restarts while an idempotent heartbeat turn is executing, the run reconciler requeues its persisted recovery envelope on the same agent_runs row at the next boot. Legacy heartbeat runs without an idempotency key or recovery envelope are marked failed (error_code = 'orphaned_restart'); the heartbeat schedule itself remains untouched.
Successful side-effecting tool steps are checkpointed durably as they finish. A fresh surviving progress checkpoint is consumed after a process restart, while a final provider stall makes the run resumable timed_out state instead of an unrecoverable failure. The scheduler uses its normal interval or provider-quota backoff before continuing from the bounded durable successful-action-key list (latest 100 entries; latest 20 shown to the model) and open todo labels. During that heartbeat continuation, an exact replay returns its stored result without executing the action again only once per checkpoint entry and only when the original result carries durable mutation or outbound-delivery evidence. Generic shell/package execution and workspace file writes (exec_command, shell_command, package_install, edit_lines, and write_file), internal agent_todo_update bookkeeping, unconfirmed mutations, failed calls, subsequent legitimate identical actions, and non-heartbeat turns are never deduplicated. Pending recovery expires after 24 hours, and a manual check cannot consume a scheduled continuation. The recovery instruction is not stored in conversation history, and successful completion clears the progress and recovery checkpoints after the final heartbeat output records recovered progress and whatever remained open.
API
GET/POST /api/heartbeat/settings?scope=user|teamchat— read/save config plus job status (lastRunAt,nextRunAt,lastDeliveryStatus:delivered/suppressed/error). Teamchat scope requires the admin role.POST /api/heartbeat/trigger— run the heartbeat immediately ({ scope }); proxies to the runtime'sPOST /internal/heartbeat/triggerand returns{ ok, delivered, suppressed, text, outputPreview, runId }.- Delivery reuses
POST /api/agent-runtime/assistant-messagewithmessageOrigin: "assistant_heartbeat".
Design notes (lessons from v1)
- Suppression is token-based, not length-based. v1 suppressed anything under 300 chars after stripping
HEARTBEAT_OK, which could silently swallow short real alerts. v2 only suppresses when the model signalsNO_MESSAGEand wrote nothing meaningful besides it, or when the output is empty. - Unchanged state is re-suppressed. The system prompt instructs the model to reply
NO_MESSAGEwhen a check finds the same situation it already posted about earlier, even if that situation is bad — a new post requires that something changed, resolved, worsened, or became newly due. - Dedicated sessions per scope. User, teamchat, and specialized-agent heartbeats never share session state, avoiding the v1 cross-scope config/session leaks.
- No stale-blocker carryover. The system prompt requires re-verification of any finding within the current run.
- Automatic delivery only. For user and teamchat scopes the agent is instructed never to post the heartbeat result through manual channel tools; the runtime delivers the final reply itself, which prevents duplicates. The agent scope inverts this: its final reply is never delivered, and the only chat output is the deliberate
team_chat_post_messagestatus posts to a configured report channel.
Provider safety refusals
A provider moderation rejection (including xAI HTTP 403 permission-denied with
SAFETY_CHECK_TYPE_BIO) fails the run with content_policy_blocked, not auth.
Heartbeat does not reconnect credentials, retry the blocked request, or replay it
on another provider. Until a task-aware safe reformulation can be verified, the
run remains failed and shows a localized refusal asking the user to review the
task. Existing side-effect checkpoints remain subject to their normal retention;
this refusal does not schedule an automatic continuation or claim completion.
The run's usage_json.runDiagnostics.error records policyCategory (e.g. bio)
and providerCode (e.g. permission-denied). Provider attempts carry the same
fields and decision: refuse_policy. Safety error messages use fixed text
instead of storing provider-echoed prompt content.
