Files
c-relay-pg/plans/cache_all_feasibility_test_plan.md

442 lines
17 KiB
Markdown

# Cache-All Feasibility Test Program
## Goal
Build a standalone C test program that connects to all outbox relays of followed
pubkeys and attempts to subscribe to **all events** (no author filter) to
determine if relays will tolerate this level of data flow. The program logs all
relay interactions (disconnections, rate limiting, CLOSED messages, NOTICE
messages) so we can assess feasibility before modifying the caching daemon.
## Background
The current caching daemon subscribes only to events from followed pubkeys
using `authors=[followed_set]` filters. The proposed "cache everything" approach
would subscribe to all events on those relays and prune old non-followed events
later. The key unknown is whether relays will allow this — they may disconnect
or rate-limit clients that request too much data.
## Preliminary Probe Results (60s on 3 major relays)
Before designing the test program, we ran 60-second probes using `nak req --stream`
with no filter (all events) against 3 major relays to measure real-world data volume
and kind distribution.
### Relay Connectivity
| Relay | Events/60s | Avg Bytes | Max Event | Unique Pubkeys | Notes |
|-------|-----------|-----------|-----------|----------------|-------|
| nos.lol | 2,634 | 1,713 | 44,372 (kind 30089) | 1,228 | Connected OK |
| relay.primal.net | 1,547 | 1,920 | 207,377 (kind 3) | 613 | Connected OK |
| relay.damus.io | 0 | — | — | — | HTTP 503 (unavailable) |
### Kind Distribution (nos.lol + primal combined, ~4,181 events)
| Kind | Description | Count | % | Avg Bytes | Notes |
|------|-------------|-------|---|-----------|-------|
| **21059** | Gift wrap (NIP-59) | 1,063 | 25.4% | 2,574 | Encrypted DMs — large, noisy |
| **5** | Deletion requests | 592 | 14.2% | 447 | Very common |
| **30078** | App-specific data | 173 | 4.1% | 3,670 | Variable size, can be large |
| **20001** | Ephemeral (key exchange?) | 253 | 6.0% | 849 | Short-lived |
| **22580** | WebRTC signaling | 193 | 4.6% | 795 | Noise |
| **1059** | Gift wrap seal (NIP-59) | 154 | 3.7% | 3,023 | Encrypted wrapper |
| **22734** | WebRTC signaling | 168 | 4.0% | 427 | Noise |
| **1** | Text notes | 58 | 1.4% | 812 | **Primary content** |
| **7** | Reactions | 59 | 1.4% | 534 | Likes |
| **0** | Profile metadata | 33 | 0.8% | 985 | Replaceable |
| **3** | Follow lists | 10 | 0.2% | 43,077 | Replaceable, can be huge |
| **10002** | Relay lists | 55 | 1.3% | 418 | Small |
| **9735** | Zap receipts | 22 | 0.5% | 2,110 | Payment confirmations |
| **6** | Reposts | 11 | 0.3% | 1,876 | |
| **30023** | Long-form articles | 3 | 0.1% | 6,119 | |
| **30089** | Chunked data | 16 | 0.4% | 43,907 | **Very large** (avg 44KB!) |
| **30001** | Lists (encrypted) | 4 | 0.1% | 30,538 | Large encrypted content |
| **30815** | Ephemeral large | 4 | 0.1% | 37,534 | **Very large** (avg 37KB!) |
| **25555** | App data | 31 | 0.7% | 5,115 | |
| **13194** | NWC info | 15 | 0.4% | 514 | Wallet connect |
| **445** | Encrypted group msg | 22 | 0.5% | 1,452 | |
| **Other** | 90+ other kinds | ~600 | ~14% | varies | Long tail of app-specific kinds |
### Key Findings
1. **Gift wrap (kind 21059) dominates**: 25% of all events are NIP-59 gift wraps
(encrypted DMs). These are large (~2.5KB avg) and mostly noise for a caching relay.
They have `expiration` tags and are ephemeral by nature.
2. **Deletion requests (kind 5) are extremely common**: 14% of all events. This is
surprising — relays are very busy processing deletes.
3. **WebRTC signaling (kinds 22580, 22734, etc.) is ~10% of traffic**: These are
ephemeral signaling events for video/voice calls. Pure noise for caching.
4. **"Social" content (kinds 1, 7, 6, 0, 3) is only ~4% of total events**: The
vast majority of relay traffic is NOT the content users actually see.
5. **Some kinds are extremely large**: Kind 30089 averages 44KB, kind 30001 averages
30KB, kind 30815 averages 37KB. These would consume significant storage.
6. **Data volume is manageable**: ~2,600-4,200 events/min across 2 relays. At this
rate, a 24h test would collect ~3.7-6M events. Storage would be ~6-10 GB of raw
JSON.
7. **Damus.io is unavailable**: Returns HTTP 503. This may be temporary or may
indicate they block unfiltered subscriptions.
### Implications for the Test Program
- **Subscribe to ALL kinds** for the test — we need to measure which kinds relays
actually send and whether they tolerate the volume.
- **Do NOT save full event JSON** during the test — at ~2MB/min, a 24h run would
produce ~3GB of raw data per relay. Instead, log kind/pubkey/size summaries.
- **Track kind 21059 (gift wrap) separately** — it's the dominant kind and may
need special handling in the real implementation.
- **Monitor for CLOSED messages** — if relays start closing subscriptions due to
volume, we'll see it in the raw relay log.
## Architecture
```
┌─────────────────────────────────────────────────────────────┐
│ cache_all_feasibility_test │
│ │
│ 1. Load root npubs from config.jsonc │
│ 2. Resolve follow graph (kind-3) → followed set │
│ 3. Discover outbox relays (kind-10002) for each follow │
│ 4. Compute minimum covering set of relays │
│ 5. Connect to all covering relays │
│ 6. Subscribe to ALL events (no authors filter) on each │
│ 7. Log every event received (counts + stats, not full JSON)│
│ 8. Log every relay message: CLOSED, NOTICE, EOSE, error │
│ 9. Run for configurable duration (default 24h) │
│ 10. Save summary report + per-relay stats to files │
└─────────────────────────────────────────────────────────────┘
```
## File Structure
All new files go in `caching/` directory, alongside the existing daemon:
```
caching/
├── Makefile # Modified: add test target
├── src/
│ ├── main.c # Existing daemon (unchanged)
│ ├── cache_all_test.c # NEW: test program entry point
│ ├── cache_all_test.h # NEW: test program header
│ ├── debug.c / debug.h # Existing (reused)
│ ├── config.c / config.h # Existing (reused for config loading)
│ ├── state.c / state.h # Existing (reused for pubkey set)
│ ├── follow_graph.c / follow_graph.h # Existing (reused)
│ ├── relay_discovery.c / relay_discovery.h # Existing (reused)
│ └── ... # Other existing files unchanged
```
## Detailed Design
### 1. Config Loading (reuse existing)
Reuse [`caching/src/config.c`](caching/src/config.c) and
[`caching/src/config.h`](caching/src/config.h) to load root npubs, upstream
relays, kinds, and follow graph settings from the same `.jsonc` config file
the daemon uses.
### 2. Follow Graph Resolution (reuse existing)
Reuse [`caching/src/follow_graph.c`](caching/src/follow_graph.c) to:
- Decode root npubs to hex
- Query each root's most recent kind-3 contact list from bootstrap relays
- Build the `cr_pubkey_set_t` of followed pubkeys
### 3. Relay Discovery (reuse existing)
Reuse [`caching/src/relay_discovery.c`](caching/src/relay_discovery.c) to:
- Query kind-10002 for each followed pubkey
- Parse "r" tags to build outbox relay map
- Compute minimum covering set of relays
### 4. Subscription Strategy
For each relay in the covering set, open a subscription with:
```c
// Filter: ALL events (no authors, no kinds, no limit)
cJSON *filter = cJSON_CreateObject();
cJSON_AddItemToObject(filter, "since", cJSON_CreateNumber((double)time(NULL)));
```
This requests every new event from `now` onward on that relay.
**Key difference from existing daemon:** No `authors` filter. This means we
receive events from *everyone* on that relay, not just followed pubkeys.
### 5. Event Handling & Logging
The `on_event` callback should:
1. **Count events per relay** (maintain a per-relay counter)
2. **Count events per kind** (maintain a kind distribution map)
3. **Count events per pubkey** (track which pubkeys are most active)
4. **Log at INFO level** every N events (e.g., every 1000) with summary stats
5. **Do NOT save full event JSON to disk during the run** (would be too much data)
6. **Periodically checkpoint** (every 5 min) a summary to a file
### 6. Relay Interaction Logging — Raw Relay Responses
**Critical requirement:** Log the **actual raw response text** from the relay,
not an interpreted summary. The relay's own words are what matter.
The `on_event`, `on_eose`, and subscription status callbacks from
[`nostr_core_lib`](../nostr_core_lib) provide status strings and message
content. These must be logged **verbatim**.
#### What to log and how:
| Trigger | Log Level | What to Log (verbatim relay text) |
|---------|-----------|-----------------------------------|
| EOSE | INFO | `[RELAY] <url> EOSE` |
| CLOSED | WARN | `[RELAY] <url> CLOSED: <relay's exact reason text>` |
| NOTICE | WARN | `[RELAY] <url> NOTICE: <relay's exact notice text>` |
| Error | ERROR | `[RELAY] <url> ERROR: <relay's exact error message>` |
| Disconnect | WARN | `[RELAY] <url> DISCONNECTED: <transport-level error if any>` |
| Reconnect | INFO | `[RELAY] <url> RECONNECTED` |
| OK (publish) | INFO | `[RELAY] <url> OK: <event_id> <relay's exact message>` |
**Do NOT** categorize or summarize. If a relay says:
```
"rate-limited: please wait 60 seconds before sending new requests"
```
log that exact string. Do NOT log just `"rate-limited"`.
#### Raw Relay Log File
In addition to the normal debug log output, write a **raw relay log file**:
```
cache_all_test_raw_relay_<timestamp>.log
```
This file contains **only** relay responses, one per line, in this format:
```
[TIMESTAMP] [RELAY] <relay_url> <RAW_RESPONSE>
```
Example:
```
[2026-08-02 13:01:00] [RELAY] wss://relay.damus.io CLOSED: "rate-limited: please wait 60 seconds before sending new requests"
[2026-08-02 13:01:05] [RELAY] wss://relay.damus.io RECONNECTED
[2026-08-02 13:01:10] [RELAY] wss://relay.damus.io NOTICE: "too many subscriptions, closing oldest"
[2026-08-02 13:02:00] [RELAY] wss://relay.primal.net EOSE
[2026-08-02 13:02:01] [RELAY] wss://relay.primal.net ERROR: "connection closed unexpectedly"
```
This file is append-only and can be `tail -f`'d during the test run.
#### Event Logging
For events received, log at TRACE level (not to the raw relay log):
```
[TRACE] [EVENT] relay=<url> kind=<N> pubkey=<first8chars>... id=<first8chars>...
```
This gives enough to correlate without flooding the log with full JSON.
Every 1000 events, log a summary at INFO level:
```
[INFO] [STATS] 5000 events received total | relayA: 3200 relayB: 1800 | kinds: 1=4500 7=500
```
### 7. Summary Report
At the end of the run (or on SIGINT/SIGTERM), write a report file:
```
cache_all_test_report_<timestamp>.txt
```
Contents:
```
=== Cache-All Feasibility Test Report ===
Duration: 24h 3m 12s
Config: ./caching_relay_config.jsonc
=== Relay Summary ===
Relay Events EOSE CLOSED NOTICE Errors Status
wss://relay.example.com 124532 12 0 2 0 OK
wss://relay2.example.com 0 0 3 5 2 BLOCKED
=== Kind Distribution ===
Kind Count %
0 1,234 0.5%
1 234,567 94.2%
3 567 0.2%
7 12,345 5.0%
9734 234 0.1%
=== Top 10 Pubkeys by Event Count ===
pubkey_hex_here... 12,345 events
pubkey_hex_here... 8,901 events
...
=== Raw Relay Response Log ===
[2026-08-02 13:01:00] wss://relay.damus.io CLOSED: "rate-limited: please wait 60 seconds before sending new requests"
[2026-08-02 13:01:05] wss://relay.damus.io RECONNECTED
[2026-08-02 13:01:10] wss://relay.damus.io NOTICE: "too many subscriptions, closing oldest"
[2026-08-02 13:02:00] wss://relay.primal.net EOSE
[2026-08-02 13:02:01] wss://relay.primal.net ERROR: "connection closed unexpectedly"
...
```
The Raw Relay Response Log section is a copy of the raw relay log file
([`cache_all_test_raw_relay_<timestamp>.log`](caching/cache_all_test_raw_relay_20260802_130000.log)).
It contains the **verbatim** text from each relay response, not interpreted
or summarized.
### 8. Per-Relay Stats File
Additionally, write a JSON file with per-relay detailed stats:
```
cache_all_test_stats_<timestamp>.json
```
This can be used for programmatic analysis. It includes the raw relay
response text for each interaction, not just counts.
## Implementation Steps
### Step 1: Create `caching/src/cache_all_test.h`
Header file declaring the test program's public interface:
```c
#ifndef CACHE_ALL_TEST_H
#define CACHE_ALL_TEST_H
/* Run the cache-all feasibility test.
* config_path: path to .jsonc config file
* duration_seconds: how long to run (0 = run until SIGINT)
* log_level: debug level 0-5
* Returns 0 on success, -1 on error.
*/
int run_cache_all_test(const char *config_path,
long duration_seconds,
int log_level);
#endif
```
### Step 2: Create `caching/src/cache_all_test.c`
Main implementation file with these sections:
1. **Includes and forward declarations**
2. **Per-relay stats tracking structure**
3. **Global stats accumulator**
4. **Callback implementations** (on_event, on_eose, on_status)
5. **Report generation** (write summary + JSON stats)
6. **Main entry point** (`run_cache_all_test`)
### Step 3: Modify `caching/Makefile`
Add a new target `cache_all_test` that compiles the test program:
```makefile
TEST_SRC = src/cache_all_test.c src/debug.c src/jsonc_strip.c src/config.c \
src/state.c src/follow_graph.c src/relay_discovery.c
cache_all_test: $(TEST_SRC) $(NOSTR_CORE_LIB)
# ... compile to ../build/cache_all_test
```
### Step 4: Build and Run
```bash
cd caching && make cache_all_test
./build/cache_all_test -c caching_relay_config.jsonc -d 3 -t 86400
```
## Output Files
| File | Contents |
|------|----------|
| `cache_all_test_raw_relay_<timestamp>.log` | **Primary output.** One line per relay response, verbatim text. Can be `tail -f`'d live. |
| `cache_all_test_report_<timestamp>.txt` | Summary report with counts, kind distribution, top pubkeys, and raw relay log section. |
| `cache_all_test_stats_<timestamp>.json` | Machine-readable JSON with per-relay stats including raw response texts. |
## What We're Measuring
1. **Relay tolerance**: Do relays CLOSE our subscription or disconnect us?
2. **Rate limiting**: How often do we get rate-limited? What are the cooldown periods?
3. **Data volume**: How many events per hour per relay? What's the kind distribution?
4. **Connection stability**: How often do relays drop us? Do they allow reconnection?
5. **Pubkey diversity**: How many unique pubkeys are posting? What's the ratio of followed vs non-followed events?
## Success Criteria
The test is considered a **success** (feasible) if:
- At least 80% of relays maintain the subscription for the full duration
- Rate limiting events are infrequent (< 5 per relay per day)
- No relay permanently bans or blacklists the connection
- Data volume is manageable (under ~1M events/day total)
The test is considered a **failure** (not feasible) if:
- Most relays CLOSE the subscription within minutes
- Rate limiting is constant (every few minutes)
- Multiple relays permanently disconnect
## Non-Goals
- This test does NOT save events to PostgreSQL
- This test does NOT modify the existing caching daemon
- This test does NOT implement pruning logic
- This test does NOT need to be efficient for production use
## Mermaid Diagram
```mermaid
flowchart TD
A[Start] --> B[Load config from .jsonc]
B --> C[Resolve follow graph kind-3]
C --> D[Discover outbox relays kind-10002]
D --> E[Compute min covering set]
E --> F[Connect to all covering relays]
F --> G[Subscribe to ALL events no authors filter]
G --> H{Test duration reached or SIGINT?}
H -->|No| I[Pump relay pool]
I --> J[Count events per relay/kind/pubkey]
J --> K[Log relay interactions CLOSED/NOTICE/errors]
K --> L[Periodic checkpoint every 5 min]
L --> H
H -->|Yes| M[Write summary report]
M --> N[Write per-relay JSON stats]
N --> O[Cleanup and exit]
```
## Files to Create
| File | Purpose |
|------|---------|
| [`caching/src/cache_all_test.h`](caching/src/cache_all_test.h) | Header with public API |
| [`caching/src/cache_all_test.c`](caching/src/cache_all_test.c) | Main implementation (~400-500 lines) |
## Files to Modify
| File | Change |
|------|--------|
| [`caching/Makefile`](caching/Makefile) | Add `cache_all_test` target |
## Files NOT Modified
The existing caching daemon files are **not touched**:
- `caching/src/main.c` — unchanged
- `caching/src/backfill.c` — unchanged
- `caching/src/live_subscriber.c` — unchanged
- `caching/src/pg_inbox.c` — unchanged
- `caching/src/forward_catchup.c` — unchanged
- `caching/src/pg_config.c` — unchanged