mirror of
https://relay.ngit.dev/npub15qydau2hjma6ngxkl2cyar74wzyjshvl65za5k5rl69264ar2exs5cyejr/ngit-grasp.git
synced 2026-10-05 15:08:24 +00:00
93 lines
3.3 KiB
Markdown
93 lines
3.3 KiB
Markdown
# Database Backend Evaluation
|
|
|
|
## Issue Summary
|
|
|
|
We currently lack guidance on when to use nostrdb vs lmdb as the database backend. We need to perform an efficiency review comparing these backends and evaluate caching strategies (especially tag caches and filter-related caching) to provide clear recommendations and optimize query performance.
|
|
|
|
## Current State
|
|
|
|
- Both nostrdb and lmdb are available as backend options
|
|
- No documented rationale for choosing one over the other
|
|
- Unknown if either implementation uses caching for common query patterns
|
|
- Unknown if tag queries are optimized with dedicated indexes/caches
|
|
- No performance benchmarks comparing the backends
|
|
|
|
## Investigation Tasks
|
|
|
|
### Backend Comparison
|
|
|
|
- [ ] Document the architecture of nostrdb backend
|
|
- [ ] Document the architecture of lmdb backend
|
|
- [ ] Identify key differences in data layout and indexing strategies
|
|
- [ ] Check if either backend has built-in caching mechanisms
|
|
- [ ] Review memory usage patterns for each backend
|
|
- [ ] Compare disk I/O characteristics
|
|
|
|
### Caching Analysis
|
|
|
|
- [ ] Identify common filter patterns in GRASP workloads (by tag, by author, by kind, etc.)
|
|
- [ ] Check if tag queries use dedicated indexes or caches
|
|
- [ ] Evaluate if frequently-used filters could benefit from caching
|
|
- [ ] Research nostrdb's native caching capabilities
|
|
- [ ] Research what caching strategies other relays use
|
|
- [ ] Determine if we need application-level caching on top of the database
|
|
|
|
### Performance Testing
|
|
|
|
- [ ] Create benchmark suite for typical GRASP queries
|
|
- Repository lookup by maintainer (author tag)
|
|
- Proposal queries (by repo, by status)
|
|
- Revision history queries
|
|
- Clone URL lookups
|
|
- [ ] Benchmark nostrdb with realistic GRASP data
|
|
- [ ] Benchmark lmdb with realistic GRASP data
|
|
- [ ] Test under concurrent read/write scenarios
|
|
- [ ] Measure query latency percentiles (p50, p95, p99)
|
|
- [ ] Measure throughput (queries/second)
|
|
- [ ] Test with varying dataset sizes
|
|
|
|
### Resource Utilization
|
|
|
|
- [ ] Compare memory footprint under load
|
|
- [ ] Compare disk space efficiency
|
|
- [ ] Evaluate startup/shutdown time
|
|
- [ ] Test database corruption recovery mechanisms
|
|
- [ ] Check if either supports hot backups
|
|
|
|
## Questions to Answer
|
|
|
|
1. **When should operators choose nostrdb vs lmdb?**
|
|
- Is one better for small deployments vs large?
|
|
- Does one scale better with dataset size?
|
|
- Are there operational trade-offs (backup, recovery, maintenance)?
|
|
|
|
2. **What caching should we implement?**
|
|
- Do we need application-level filter result caching?
|
|
- Should we cache tag lookups separately?
|
|
- What cache invalidation strategy makes sense?
|
|
|
|
3. **What are the performance characteristics?**
|
|
- Query latency for common operations
|
|
- Write throughput for event ingestion
|
|
- Read throughput for serving subscriptions
|
|
|
|
4. **What are the operational considerations?**
|
|
- Ease of debugging/inspection
|
|
- Backup and restore procedures
|
|
- Database migration strategies
|
|
|
|
## Findings
|
|
|
|
*(To be filled in during investigation)*
|
|
|
|
## Implementation Plan
|
|
|
|
*(To be determined after investigation)*
|
|
|
|
## Success Criteria
|
|
|
|
- Clear documentation on when to use each backend
|
|
- Performance benchmarks published for both backends
|
|
- Identified and implemented any beneficial caching strategies
|
|
- Recommendations added to deployment documentation
|