Trace profiling on Pixel 8 revealed two massive overhead sources:
1. ExternalSyntheticBackport0.m — 17.5% of verify (3.45ms)
D8/R8 desugaring wraps Math.unsignedMultiplyHigh in a synthetic
backport because minSdk=26 < 35. The backport adds 139ns per call
× 24,875 calls. The actual UMULH intrinsic is ~1ns, but the
wrapper adds ~138ns.
FIX: Use pure-Kotlin unsignedMultiplyHighFallback on Android instead
of Math.unsignedMultiplyHigh. The fallback (4 Long multiplies +
shifts, ~10-20ns) is FASTER than the backported intrinsic (139ns).
Removed all API-level dispatch — a single fallback path for all
Android versions.
2. uLt function calls — 20.7% of verify (4.08ms)
The expect/actual uLt (can't be inline) was called 49,924 times
per verify at 82ns each. Most calls came from the fused
fieldMulReduceWith/fieldSqrReduceWith inline expansions.
FIX: Add uLtInline — a private inline function using the XOR trick
directly. Since it's @Suppress("NOTHING_TO_INLINE") inline, the
Kotlin compiler inlines it at every call site (not the JIT).
Replaces all 55 uLt calls in the fused inline functions.
JVM is unaffected (uses unfused path with Long.compareUnsigned).
Combined: eliminates ~38% of verify overhead on Android.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
BIP-340 public keys always have even y, so the y-parity prefix byte
(0x02) is redundant when the caller already has the 32-byte x-only
pubkey. This new overload takes the x-only pubkey directly, avoiding:
1. The expensive G multiplication to derive the pubkey (~20μs on Android)
2. The 33→32 byte array copy that signSchnorrWithPubKey does internally
Added to:
- Secp256k1.signSchnorrWithXOnlyPubKey (core implementation)
- Secp256k1Instance (expect/actual, falls back to C lib's signSchnorr
since the native C lib always derives pubkey internally)
- Secp256k1InstanceOurs (uses the optimized pure-Kotlin path)
- Nip01Crypto.signWithPubKey (app-level convenience)
The app's KeyPair already stores the 32-byte x-only pubkey. Callers
like EventAssembler.hashAndSign can pass it through to skip the
G multiplication entirely.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
Document WHY each approach was chosen, what alternatives were tested,
and what the measured impact was — so future contributors don't
accidentally revert optimizations or repeat failed experiments.
Key decisions documented:
- uLt() expect/actual: why XOR on Android, Long.compareUnsigned on JVM,
and why a shared inline fun in commonMain caused 30% JVM regression
- fieldMulReduceWith: why fused mul+reduce, why inline+crossinline,
why NOT 5x52 limbs, why NOT single-method with all API branches
- @JvmField: why it's required on MutablePoint/AffinePoint/PointScratch,
with bytecode counts showing ~7,450 + ~2,000 virtual getter calls
eliminated per verify
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
The previous approach inlined all 3 API-level branches into a single
fieldMulReduce method (~600 DEX instructions). ART's register allocator
produces suboptimal code for such large methods, causing stack spills
that negate the benefit of eliminating wrapper calls.
Split each API-level path into its own private function (~200 DEX
instructions each). ART JIT-compiles only the hot one (fieldMulApi35
on API 35+) with full optimization, while the tiny dispatch function
(fieldMulReduce) is easily devirtualized and inlined.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
On Android (ART), the unsignedMultiplyHigh wrapper function has per-call
branching for API level detection that prevents ART's JIT from inlining
the intrinsic into the hot loop. This creates ~10,000 extra function
calls per signature verify (20 wrapper calls × 500 field muls).
This commit introduces fieldMulReduce/fieldSqrReduce as expect/actual
functions that fuse U256.mulWide + FieldP.reduceWide into a single
compilation unit. The multiply-high intrinsic is passed as a crossinline
lambda and inlined at each call site, producing platform-specific code
with zero wrapper overhead:
- Android API 35+: Math.unsignedMultiplyHigh inlined directly (UMULH)
- Android API 31-34: Math.multiplyHigh + correction inlined (SMULH)
- Android API <31: pure-Kotlin fallback inlined
- JVM: Math.unsignedMultiplyHigh inlined directly
- Native: pure-Kotlin fallback inlined
The API level check happens ONCE per fieldMulReduce call (outermost
branch) rather than per multiply-high call (innermost loop), so ART
profiles and JIT-compiles only the hot path.
https://claude.ai/code/session_01EMY5RnXb9rnsyU2KbXrSaY
ThreadLocal.withInitial is a JVM-only API that doesn't compile on
Kotlin/Native or iOS targets. Replace all usages in commonMain
(FieldP.kt, Point.kt) with a new ScratchLocal expect/actual:
- jvmMain/androidMain: delegates to java.lang.ThreadLocal (same behavior)
- nativeMain: holds value directly (Kotlin/Native coroutines are
cooperative, scratch buffers don't need thread isolation)
https://claude.ai/code/session_01BhU63WUe9AhikZxRdw3Lpg
Three Android-specific optimizations:
1. Use Math.unsignedMultiplyHigh on API 35+ (Android 15): single UMULH
instruction, eliminates the 4-insn signed→unsigned correction that
the fallback path requires. Same optimization as our JVM 18+ path.
2. Use Math.multiplyHigh + correction on API 31-34 (Android 12-14):
avoids the pure-Kotlin 4×32-bit sub-product fallback entirely.
3. Resolve API level check ONCE at class load via static final fields
(HAS_MULTIPLY_HIGH, HAS_UNSIGNED_MULTIPLY_HIGH) instead of checking
Build.VERSION.SDK_INT on every call. These functions are called 16×
per field multiply (~12,000× per signature verify), so eliminating
the per-call branch matters.
Performance tiers on Android:
API 35+ (Android 15): ~same as JVM 18+ (UMULH intrinsic)
API 31-34 (Android 12-14): SMULH + 3 correction insns per product
API 26-30 (Android 8-11): pure-Kotlin fallback (4 sub-products)
https://claude.ai/code/session_01BhU63WUe9AhikZxRdw3Lpg
Foundation layer for the 4×64-bit limb optimization. This commit rewrites
U256 from IntArray(8) to LongArray(4) and adds platform-specific
Math.multiplyHigh for 64×64→128-bit products:
- JVM: Math.multiplyHigh intrinsic (single IMULH/SMULH instruction)
- Android API 31+: Math.multiplyHigh, fallback to pure Kotlin on older
- Native: pure Kotlin fallback (4 sub-products per multiplyHigh call)
The 4×64 representation reduces inner products from 64 to 16 per field
multiply on JVM, a potential ~1.5-2× speedup on the critical path.
NOTE: This commit intentionally breaks the build — FieldP, ScalarN, Glv,
Point, KeyCodec, Secp256k1, and all tests still expect IntArray(8).
They will be updated in subsequent commits.
https://claude.ai/code/session_01BhU63WUe9AhikZxRdw3Lpg
Implements all secp256k1 operations used by Secp256k1Instance in pure Kotlin,
eliminating the dependency on native secp256k1 bindings (fr.acinq.secp256k1-kmp).
Implementation:
- Field.kt: 256-bit unsigned integer arithmetic, field mod p, scalar mod n
- Point.kt: EC point operations (Jacobian coords), point parsing/serialization
- Secp256k1.kt: Public API - pubkeyCreate, pubKeyCompress, secKeyVerify,
signSchnorr/verifySchnorr (BIP-340), privKeyTweakAdd, pubKeyTweakMul
Tests (36 total):
- All 19 BIP-340 test vectors (signing vectors 0-3, 15-18; verify vectors 4-14)
- ACINQ test vectors for key creation, compression, secKeyVerify, privKeyTweakAdd,
pubKeyTweakMul
- ECDH symmetry and sign/verify round-trip tests
Secp256k1Instance is now a concrete object in commonMain instead of expect/actual,
delegating to the pure-Kotlin implementation. All existing NIP-44 and BIP-32 tests
pass with the new implementation.
https://claude.ai/code/session_01BhU63WUe9AhikZxRdw3Lpg
Before (4 steps, manual wiring):
val transport = AndroidBleTransport(context)
val mesh = BleMeshManager(transport, object : BleMeshListener { ... })
transport.setListener(mesh)
mesh.start()
After (2 steps, auto-wired):
val mesh = BleNostrMesh(AndroidBleTransport(context))
mesh.onEvent { event, peer -> saveEvent(event) }
mesh.start()
The facade uses lambda callbacks instead of interface implementation,
and auto-wires the transport listener via AndroidBleTransportContract.
BleMeshManager remains available for advanced use cases.
https://claude.ai/code/session_01Tz5E73Rj7tL48A3qUGS5DT
Reviewed the samiz BLE implementation and fixed compatibility issues:
- Chunk index: changed from 2 bytes to 1 byte (matching samiz/NIP-BE example)
- Chunk overhead: 2 bytes total (1 index + 1 total count), not 3
- chunkSize parameter now means payload size (500), not total chunk size
- Android: TX power HIGH, advertise timeout indefinite (matching samiz)
- Android: request MTU 512 and CONNECTION_PRIORITY_HIGH on connect
- Android: discover services after MTU negotiation (samiz flow)
- Android: set WRITE_TYPE_DEFAULT on write characteristic
- MTU-to-chunkSize: properly accounts for ATT overhead (3) + chunk overhead (2)
These changes ensure wire-level compatibility with samiz devices.
https://claude.ai/code/session_01Tz5E73Rj7tL48A3qUGS5DT
Adds full NIP-BE implementation for BLE-based Nostr peer-to-peer
messaging and synchronization with KMP support across all targets.
Protocol layer (commonMain):
- BleConfig: NIP-BE constants (service UUID, characteristic UUIDs)
- BleMessageChunker: DEFLATE compression + chunk splitting/joining
- BleRole/assignRole: Role assignment based on UUID comparison
- BleTransport: Platform-agnostic BLE transport interface
- BleNostrClient: IRelayClient implementation over BLE
- BleNostrServer: Relay server over BLE with notification support
- BleMeshManager: High-level mesh manager with auto-discovery,
role assignment, and event broadcasting
- BleChunkAssembler: Thread-safe chunk reassembly
Platform support:
- Deflate expect/actual for jvmAndroid, apple, and linux targets
- AndroidBleTransport: Full Android BLE implementation using GATT
server/client APIs with scanning, advertising, and MTU negotiation
Tests: Deflate compression, message chunking, role assignment,
and chunk assembly (25 tests passing).
https://claude.ai/code/session_01Tz5E73Rj7tL48A3qUGS5DT
- Add optional throwable parameter to d() and i() across all platforms
- Add inline lambda overloads for all log levels to defer message
construction, avoiding string allocation when level is filtered
- Restore VoiceMessagePreview to pass throwable directly
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Deprecated APIs in Quartz are kept for library consumers but internal usages
now carry @Suppress annotations so the module builds warning-free. Also
replaces redundant .toInt() on hex Int literals in ChaCha20Core and migrates
deprecated java.net.URL to URI.resolve() in ServerInfoParser.
The spotless license header delimiter is updated to recognize @file: annotations.
https://claude.ai/code/session_01EwS56YAGGnnac5EuwhaUSs