Files
sovereign_browser/plans/processes-window.md
T

296 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Processes Window — Implementation Plan
## Goal
A `sovereign://processes` internal page that lets the user pinpoint which
open tab is consuming CPU and dig into *what that page is doing*. Primary
use case: diagnosing `~/lt/client`, which currently burns CPU across many
tabs.
## The core technical reality
The codebase deliberately **shares one WebProcess across tabs** — every
new tab/webview is created with `webkit_web_view_new_with_related_view()`
(see [`tab_manager.c`](../src/tab_manager.c:135) and the call sites at
lines 1106, 1228, 2657, 3602, 3624). This was done to avoid a
`std::optional<WindowFeatures>` assertion crash in WebKitGTK.
Consequence: process-level CPU% from `/proc` tells you which *WebProcess*
is hot, but **not which tab within it**. If 10 `~/lt/client` tabs share
one WebProcess at 90% CPU, `/proc` alone cannot attribute the load.
The plan therefore uses **two layers**:
1. **Layer 1 — Process view** (from `/proc`): main, WebKitWebProcess(es)
with their hosted tabs listed, WebKitNetworkProcess, Tor, FIPS. Catches
the single-tab-per-process case and shows overall footprint.
2. **Layer 2 — Per-tab page probe** (injected JS in every webview): this
is what actually pinpoints the offending tab *within* a shared
WebProcess and tells you what it is doing.
## Architecture
```mermaid
flowchart LR
subgraph BrowserProcess[sovereign_browser main PID]
ProcInfo[process_info.c reads /proc]
Bridge[nostr_bridge.c sovereign:// routes]
AgentServer[agent_server / MCP]
end
subgraph WebProcess[WebKitWebProcess pid N - shared by many tabs]
ProbeA[perf-probe.js in tab A]
ProbeB[perf-probe.js in tab B]
ProbeC[perf-probe.js in tab C]
end
ProbeA -->|sovereign://processes/probe-report POST| Bridge
ProbeB -->|sovereign://processes/probe-report POST| Bridge
ProbeC -->|sovereign://processes/probe-report POST| Bridge
ProcInfo -->|enumerates /proc and webkit_web_view_get_process_id| Bridge
Bridge -->|JSON| UI[www/processes.js renders tables]
Bridge -->|JSON| AgentServer
```
## Layer 1 — Process view (`src/process_info.{c,h}`)
### Process discovery
| Process | How found |
|---|---|
| Main (`sovereign_browser`) | `getpid()` |
| WebKitWebProcess (≥1) | For each tab, `webkit_web_view_get_process_id(tab->webview)` returns the WebProcess PID. Group tabs by PID. |
| WebKitNetworkProcess | Scan `/proc/*/comm` for `WebKitNetworkProcess` whose PPID is the main PID (read `/proc/<pid>/stat` field 4). |
| WebKitGPUProcess / StorageProcess | Same PPID scan, if present. |
| Tor (managed) | `net_service_get_status(NET_SERVICE_TOR)->pid` when `ownership == OWNERSHIP_MANAGED`. |
| FIPS (managed) | `net_service_get_status(NET_SERVICE_FIPS)->pid` when `ownership == OWNERSHIP_MANAGED`. |
### Per-process fields (all from `/proc`, no external deps)
| Field | Source |
|---|---|
| pid | — |
| name | `/proc/<pid>/comm` |
| cmdline | `/proc/<pid>/cmdline` (NUL-split, space-joined) |
| state | `/proc/<pid>/stat` field 3 (R/S/D/Z/T) |
| ppid | `/proc/<pid>/stat` field 4 |
| cpu_percent | delta of (`utime`+`stime`, fields 14+15) between two reads, divided by elapsed wall time × `sysconf(_SC_CLK_TCK)` |
| rss_kb | `/proc/<pid>/status``VmRSS` |
| pss_kb | `/proc/<pid>/smaps_rollup``Pss` (fallback to RSS if unavailable) |
| uss_kb | `/proc/<pid>/smaps_rollup``Private_Clean` + `Private_Dirty` |
| vmpeak_kb | `/proc/<pid>/status``VmPeak` |
| vmswap_kb | `/proc/<pid>/status``VmSwap` |
| threads | count of entries in `/proc/<pid>/task/` |
| uptime_sec | `time(NULL) - (btime + starttime/clk_tck)`; btime from `/proc/stat` |
| io_read_kb | `/proc/<pid>/io``read_bytes` / 1024 |
| io_write_kb | `/proc/<pid>/io``write_bytes` / 1024 |
| fd_count | count of entries in `/proc/<pid>/fd/` |
| ownership | `main` / `webkit-renderer` / `webkit-network` / `webkit-gpu` / `tor` / `fips` |
| service_state | for Tor/FIPS: from `net_service_t.state` (READY/BOOTSTRAPPING/…) |
| hosted_tabs | for renderers: array of `{index, title, url}` from `tab_manager` |
### CPU% computation
Keep a static `GHashTable<pid, prev_sample>` between calls. Each call:
1. Read current `utime+stime` and `time(NULL)`.
2. `cpu% = (cur - prev) / clk_tck / (now_wall - prev_wall) * 100`.
3. Store current as prev.
Caller polls `sovereign://processes/list` every 1s; first call returns
0% (no baseline), second call onward is accurate. Same approach as
`top`/`htop`.
### API
```c
/* src/process_info.h */
cJSON *process_info_get_processes_json(void); /* Layer 1 table */
cJSON *process_info_get_tabs_json(void); /* tab→WebProcess mapping + last probe */
cJSON *process_info_get_tab_probe_json(int tab_index); /* drill-down for one tab */
void process_info_record_probe(int tab_index, cJSON *probe); /* called from probe-report route */
```
## Layer 2 — Per-tab page probe (`www/js/perf-probe.js`)
Injected into every webview via `WebKitUserContentManager` (same pattern
as the existing `window.nostr` injection in
[`nostr_inject.c`](../src/nostr_inject.c:1)). Runs before page scripts
where possible (using `WebKitUserContentInjectedFramesAllFrames` and
`WebKitUserScriptInjectAtDocumentStart`).
### Signals collected
| Signal | Source | What it tells you |
|---|---|---|
| **long_tasks** | `PerformanceObserver({entryTypes:['longtask']})` | Every main-thread task >50ms with duration + attribution. *Direct answer to "which tab is hogging CPU".* |
| **cpu_busy_percent** | sum of long-task durations in last 1s window | True per-tab CPU-busy %. |
| **fps** | `requestAnimationFrame` cadence over 1s | Constant repainting / animation loops. |
| **timer_count** | patched `setTimeout`/`setInterval` (count active) | Runaway polling loops (common in dashboards like `~/lt/client`). |
| **timer_top** | top 5 intervals by frequency in last 1s | Which polling endpoints/scripts. |
| **net_requests** | `PerformanceObserver({entryTypes:['resource']})` | Fetch/XHR/WebSocket storms, polling endpoints. |
| **net_in_flight** | patched `fetch`/`XMLHttpRequest`/`WebSocket` | Current open requests. |
| **heap_used_mb** | `performance.memory.usedJSHeapSize` (if exposed) | Memory leaks, growing arrays. |
| **event_listener_count** | patched `addEventListener` | Leaking listeners. |
| **worker_count** | patched `Worker` constructor | Background CPU not visible as long tasks. |
| **dom_node_count** | `document.getElementsByTagName('*').length` | Page bloat. |
| **long_task_timeline** | ring buffer of last 30s of long tasks | Drill-down chart. |
| **long_task_top_sources** | aggregate by `entry.name` / script URL | Which script is responsible. |
### Reporting
The probe batches a report every 1s and POSTs it to
`sovereign://processes/probe-report?tab_index=N` with JSON body. The
route handler in `nostr_bridge.c` calls
`process_info_record_probe(N, body)`, which stores the latest probe per
tab index in a static array (sized to `tab_manager_count()`).
The probe must NOT run on `sovereign://` internal pages (settings, fips,
processes itself, agents) — skip injection when the URL scheme is
`sovereign://`. This avoids self-noise and recursion.
### Probe payload shape
```json
{
"tab_index": 3,
"cpu_busy_percent": 42.5,
"fps": 12,
"timer_count": 38,
"timer_top": [{"code":"pollStatus","interval_ms":250,"count":4}],
"net_in_flight": 3,
"net_requests_last_sec": 8,
"net_top_endpoints": [{"url":"wss://.../events","count":4}],
"heap_used_mb": 124.3,
"event_listener_count": 217,
"worker_count": 2,
"dom_node_count": 4821,
"long_tasks_last_sec": [{"duration_ms":180,"name":"setTimeout","attribution":"app.js:420"}],
"long_task_top_sources": [{"name":"app.js","total_ms":1240,"count":8}]
}
```
## Routes (added to `nostr_bridge.c`)
Following the [`sovereign://fips`](../src/nostr_bridge.c:4227) pattern:
| Route | Handler | Returns |
|---|---|---|
| `sovereign://processes` | `serve_embedded_file("processes.html")` | HTML |
| `sovereign://processes.css` | `serve_embedded_file("processes.css")` | CSS |
| `sovereign://processes.js` | `serve_embedded_file("processes.js")` | JS |
| `sovereign://processes/list` | `process_info_get_processes_json()` | JSON, polled 1s |
| `sovereign://processes/tabs` | `process_info_get_tabs_json()` | JSON, polled 1s |
| `sovereign://processes/tab_probe?index=N` | `process_info_get_tab_probe_json(N)` | JSON, on drill-down |
| `sovereign://processes/probe-report?tab_index=N` | `process_info_record_probe(N, body)` | 204 No Content |
## UI (`www/processes.{html,css,js}`)
Follows the [`www/fips.*`](../www/fips.html:1) tabbed pattern.
### Tab 1 — Processes
Sortable table, default sort by `cpu_percent` desc:
| PID | Name | CPU% | RSS | PSS | Threads | Uptime | State | Tabs |
|---|---|---|---|---|---|---|---|---|
| 12345 | sovereign_browser | 2.1 | 180M | 160M | 12 | 1h23m | R | — |
| 12367 | WebKitWebProcess | 88.4 | 820M | 740M | 8 | 1h22m | R | 3,7,9,12,15 |
| 12368 | WebKitNetworkProcess | 1.2 | 90M | 80M | 4 | 1h22m | S | — |
| 12400 | tor | 0.3 | 45M | 40M | 3 | 1h20m | S | — |
Click a WebKitWebProcess row → expands inline to list its hosted tabs
with their Layer-2 `cpu_busy_percent` and `fps` so you can see the
hot tab even within a shared process.
### Tab 2 — Tabs (the primary diagnostic view)
Sortable table, default sort by `cpu_busy_percent` desc:
| # | Title | URL | CPU-busy% | FPS | Timers | Net | Heap | DOM | WPID |
|---|---|---|---|---|---|---|---|---|---|
| 9 | Client — Dashboard | https://client/... | 41.2 | 11 | 38 | 3 | 124M | 4821 | 12367 |
| 7 | Client — Alerts | https://client/... | 0.8 | 60 | 2 | 0 | 22M | 410 | 12367 |
| 3 | sovereign://settings | — | — | — | — | — | — | — | 12367 |
Click a tab row → **drill-down panel** opens below with:
- Long-task timeline (last 30s, sparkline)
- Top long-task sources (script URLs + total ms)
- Active timers list (code label, interval, count)
- Recent network requests (URL, type, count)
- Worker count, event listener count, heap trend
- Action buttons: **Reload tab**, **Suspend tab** (replace with
`about:blank` keeping URL), **Close tab**
### Styling
Reuse `sovereign-base.css` and the card/table styles from
[`www/fips.css`](../www/fips.css:1). Add a heat-bar column background
(green→yellow→red by CPU%) for instant visual scan.
## Menu integration
- Hamburger menu item **"Processes…"** in [`main.c`](../src/main.c:325)
next to "FIPS Mesh…", navigates active tab to `sovereign://processes`
(or opens a new tab if none active) — same pattern as
`open_fips_cb` at line 332.
- Keyboard shortcut `Ctrl+Shift+Esc` (matches Windows Task Manager
convention) registered in [`shortcuts.c`](../src/shortcuts.c:78) as
`open_processes`.
## MCP tools (so an agent can self-diagnose)
Add to [`agent_tools.c`](../src/agent_tools.c:1) /
[`agent_mcp.c`](../src/agent_mcp.c:1):
| Tool | Args | Returns |
|---|---|---|
| `processes.list` | none | Layer 1 JSON (all PIDs + fields) |
| `processes.tabs` | none | Layer 2 JSON (per-tab probes) |
| `processes.tab_probe` | `{tab_index}` | drill-down JSON for one tab |
This lets you ask the agent "find the tab burning CPU in `~/lt/client`
and tell me what it's doing" — the agent calls `processes.tabs`, finds
the high `cpu_busy_percent` tab, calls `processes.tab_probe` for the
drill-down, and reports the offending script/timer/endpoint.
## Build / packaging
- Add `src/process_info.c` to `Makefile` sources.
- Add `www/processes.{html,css,js}` and `www/js/perf-probe.js` to
[`embed_web_files.sh`](../embed_web_files.sh:1) so they are embedded
into the binary via `serve_embedded_file()`.
- Register the perf-probe user script in the shared
`WebKitWebContext`'s `WebKitUserContentManager` at startup (in
`main.c` near where `nostr_inject` is wired), with a URL filter that
excludes `sovereign://*`.
## Verification
1. `make`
2. `./browser.sh restart --login-method generate`
3. Open `sovereign://processes` — verify main + WebKit processes listed
with sane CPU%/RSS.
4. Open several tabs of a CPU-heavy page (e.g.
[`tests/local-site/media.html`](../tests/local-site/media.html:1) or
a synthetic `<script>while(true){}</script>` page) — verify the Tabs
view ranks them by `cpu_busy_percent` and the drill-down shows
long-task sources.
5. Open `~/lt/client` across multiple tabs — verify the hot tab is
identifiable and the drill-down shows which script/timer/endpoint is
responsible.
6. Via MCP: call `processes.tabs` and `processes.tab_probe` and verify
JSON shape matches the contract above.
## Out of scope (future work)
- Forcing process-per-tab via `WebKitWebsitePolicies` /
`WEBKIT_PROCESS_MODEL_MULTIPLE_SECONDARY_PROCESSES` — would give true
per-tab `/proc` attribution but breaks the existing
`related_view` sharing workaround. Tracked separately.
- Historical recording (save probe samples to SQLite for trend charts).
- Per-tab network throttling / CPU limiting (would need WebKit API
support that may not exist).
- Killing/suspending a tab's WebProcess specifically (currently only
per-tab reload/close, which is safe).