Skip to content

How much faster is Aster than Mihomo? ​

The short version ​

Aster does not magically make your network faster. It removes a lot of repeated small work inside the proxy core. The medians below come from three identical rounds of the same program and tests on a real OpenWrt soft router:

Workload you actually hitHow much fasterWhy
Moving 32 KiB over TCPabout 2% more throughput, about 2% less processing timeReuse the scratch buffer that moves data
Packing a 1 KiB AnyTLS frameabout 2.0× fasterWrite straight into a reusable buffer; no extra object wrapper
Packing a 16 KiB AnyTLS frameabout 1.4× fasterSame as above, plus fewer small allocations
Preparing one UDP packetabout 5.4× fasterRecycle the packet data structure and reuse it next time
Ignoring one disabled debug logabout 101× fasterReturn before the string or log event is built

IMPORTANT

The closest “actually moving data” number is the TCP improvement of about 2%. 5.4× and 101× are tiny core steps that run many times. They do not mean download speed becomes 5.4× or 101×.

NOTE

The latest 2026-09-07 Linux WSL2 A/B compares Aster before and after PR #4, not Aster against Mihomo. Median RSS with 4,096 DNS cache entries fell 2.5%, with overlapping ranges; 1,000 TCP connections did not use less RAM. Large-pool operations and full DNS-message cache hits became slower. The results below include these costs, not a download-speed claim.

What actually changed? ​

1. Stop allocating new memory for every packet ​

TCP, UDP, and AnyTLS used to create new scratch objects for each piece of data. Those objects are now recycled and reused. The CPU spends less time on garbage, and busy loads are less likely to hitch.

2. Connections do not fight over one lock ​

Looking up rules or a proxy used to read shared data under a lock. Now a complete snapshot is published when configuration updates. Packet processing only reads the snapshot and does not queue on one another.

3. UDP no longer repeats the same conversions ​

Source addresses stay in a form the computer can compare directly, instead of being converted to a string on every packet. Socket idle timers are refreshed in batches instead of being reset on every received packet.

4. AnyTLS rules are computed once ​

Padding rules are parsed when the configuration loads. When data is sent, Aster looks up the result and builds the frame directly. It no longer splits strings and recalculates while transmitting.

5. Logs that will not be shown do no work ​

If debug logging is off and no dashboard is listening, Aster returns before formatting the message. Previously it still built a full log and then discarded it.

6. Traffic numbers increment in place ​

Upload, download, and connection counts use cheap incremental stats. Each traffic update only changes the numbers. It no longer rescans every active connection to get a total.

Full test data ​

2026-09-07 PR #4: Linux regression and memory A/B ​

Baseline be6d7912 → current core c5f553dc (PR #4: bounded allocator slabs and compact A/AAAA cache). This is Aster versus Aster, not Mihomo, and does not mean the subsequent memory optimization plan has been implemented.

Environment and method ​

  • Arch Linux under WSL2, kernel 6.6.114.1-microsoft-standard-WSL2; Ryzen 9 5900X, 24 logical CPUs, approximately 31.3 GiB available to WSL. This is not OpenWrt or an isolated, fixed-frequency physical Linux host.
  • Both revisions use Linux go1.26.5-X:nodwarf5, GOAMD64=v1, CGO_ENABLED=0, and -trimpath -ldflags='-s -w'. Whole-process GOMAXPROCS=2; microbenchmarks use 1. No GOGC/GOMEMLIMIT settings or forced GC.
  • Four scenarios, seven rounds each, fresh processes, alternating before/after order: 56 complete core processes. Each scenario has identical configuration SHA256 hashes across revisions.
  • Rule mode, MATCH,DIRECT, TUN off, loopback SOCKS/controller. DNS and IPv6 are off except in the DNS scenario. Debug is enabled for the controller, without a log WebSocket subscriber.
  • After establishing the workload, wait 15 seconds and sample every 250 ms for three seconds. Report the median of seven per-process medians and the complete range. Read VmRSS, VmHWM, and VmSwap from /proc/<pid>/status; every sampled swap value is zero. These are not Windows working-set measurements and exclude some kernel socket/splice-pipe memory.
  • TCP uses SOCKS to an external Python loopback echo server. Every connection verifies echo and controller counts, then transfers 128 × 4096 bytes in each direction, with at most 16 workers. After closing all connections and confirming a zero controller count, wait 60 seconds. Clients bind 127.0.0.2 to avoid unrelated local TCP four-tuples sharing a source AddrPort and triggering loop detection.
  • DNS uses LRU capacity 8192 and 4096 distinct names, half A and half AAAA, TTL 3600. Every process verifies exactly 4096 upstream queries during priming and zero on a second pass. TTLs and cache capacity are not reduced to improve results.

Complete-process RSS ​

MiB; parentheses show the complete seven-run range. Raw data also includes each phase's VmHWM and 12 samples.

ScenarioBefore RSS (range)After RSS (range)Median change
Idle35.11 (34.30–38.80)34.82 (34.13–38.80)−0.8%
Hold 100 TCP connections36.05 (35.76–36.34)35.71 (35.42–36.05)−1.0%
100 TCP connections, 60 s after release36.28 (35.88–36.34)35.85 (35.42–36.19)−1.2%
Hold 1,000 TCP connections58.35 (57.32–61.21)60.93 (57.49–61.90)+4.4%
1,000 TCP connections, 60 s after release59.42 (58.52–64.46)63.93 (58.69–65.01)+7.6%
4,096 A/AAAA DNS cache entries44.19 (43.75–47.04)43.09 (42.63–45.90)−2.5%

Every before/after range overlaps. DNS's median is approximately 1.10 MiB lower, not a guaranteed saving. Median RSS for 1,000 TCP connections increased. Neither revision returned to idle after release. Post-transfer, still-open TCP median RSS was unchanged from the held stage; see evidence for all ranges. This run does not establish that bounded pools return whole-process RAM to the OS within 60 seconds.

CPU/allocation costs: not a universal speedup ​

Each case/sample starts a new go test -c binary process, with identical benchmark functions across revisions: two seconds per case, -test.cpu=1, seven interleaved rounds, 112 microbenchmark processes. Medians below; all samples, ranges, and benchstat output are in the evidence archive.

OperationBefore → AfterB/op; allocs/op: before → after
Allocator 8 KiB Get/Put15.35 → 15.29 ns (not significant)0; 0 → 0; 0
Allocator 16 KiB Get/Put15.72 → 36.66 ns (+133.2%)0; 0 → 0; 0
Allocator 32 KiB Get/Put15.59 → 37.28 ns (+139.1%)0; 0 → 0; 0
Allocator 128 KiB Get/Put15.38 → 37.32 ns (+142.7%)0; 0 → 0; 0
DNS GetMsgFromCacheHit199.1 → 397.6 ns (+99.7%)252; 5 → 516; 10
DNS ExchangeContextCacheHit345.0 → 529.0 ns (+53.3%)252; 5 → 516; 10
DNS LookupIPv4CacheHit200.9 → 203.8 ns (not significant)88; 3 → 88; 3
Relay 32 KiB4.302 → 4.252 µs (not significant)0; 0 → 0; 0

Large channel pools trade higher Get/Put costs for bounded retention. Full DNS-message reads expand compact entries; the measured time and allocation regressions are real costs of this revision. The five timing regressions above have benchstat p=0.001, without multiple-comparison correction. IPv4 lookup has p=0.805, relay p=0.902; neither is claimed faster. This reports the merged code without silently changing the core or hiding regressions.

Heap diagnostic limitations ​

Six additional diagnostic processes capture before/after inuse_space for idle, 1,000 TCP connections, and DNS using /debug/pprof/heap?gc=0. Profiles are collected outside the RSS cohort to avoid contaminating later measurement phases.

  • DNS before flat top three: protobuf RegisterFile.func2, 514.38 KiB; regexp compiler, 513.50 KiB; dns.cloneMsg, 512.07 KiB. After: gopacket initialization, 525.43 KiB; protobuf type registry, 513.12 KiB; dns.compactIPs, 512.04 KiB.
  • TCP before has only two nonzero flat entries: regexp2 initialization, 518.65 KiB, and DNS map initialization, 512.88 KiB. After: regexp.onePassCopy, 521.05 KiB, and sync.OnceValue, 512.01 KiB. There is no third nonzero entry to report, and no defensible per-connection heap-byte estimate.
  • With natural GC and default sampling, heap profiles may lag current allocations. The TCP profiles still primarily describe initialization, so they do not satisfy the plan's P0-1 attribution acceptance criteria. Raw profiles/top output are retained rather than forcing GC for cleaner-looking numbers.

Verification and evidence ​

  • Windows Go 1.26.3: complete normal/with_low_memory tests and builds, pool/DNS race tests in both modes, go vet ./..., changed-Go-file gofmt checks, and golangci-lint (zero issues) passed.
  • Linux Go 1.26.5: go test -p=4 -count=1 -timeout=10m ./... and the complete low-memory suite passed, as did pool/DNS race tests in both modes, go vet ./..., and both builds. The initial three-minute timeout is retained; listener/inbound took 209.051 seconds on retry, with no tests changed or skipped to pass.
  • Early probes encountered loopback source-port collisions and AAAA suppression by the global IPv6 setting. The harness was corrected and affected scenarios rerun. Failed probes are retained in pilot/, not mixed into the final 56 processes.
  • Download evidence, harnesses, benchstat, and reproduction instructions. No executables. The core was not modified; documentation commits do not change the measured core hashes.
  • Linux TLS-proxy/WAN, OpenWrt, UDP queue pressure, low-memory performance, and full-process allocator-burst retention A/B were not measured. This is not acceptance of every subsequent optimization-plan item.

2026-09-05 process memory and core-cost A/B ​

The baseline is fixed at 4a59a634: previously shipped gains are not counted again. Of 24 Paseo-managed Pi Grok 4.6 XHigh investigations, 21 approved implementations were integrated after Astra review. Routing without a safe win, unsupported global pool-class changes, and dial cancellation with unresolved winner/TFO lifetime risks were deferred. Fast command delivery was verified, but server-side priority was not independently observed.

Whole processes: not B/op, and not Linux RSS ​

The final comparison is 4a59a634 → 399e7f83. Windows 11, Ryzen 9 5900X, 64 GB DDR4, Go 1.26.3, GOAMD64=v1, CGO_ENABLED=0, and -trimpath -ldflags='-s -w'. Runtime GOMAXPROCS=2; GOGC/GOMEMLIMIT unset; no forced GC. Builds and agent work stopped before serial measurements. Six scenarios × seven alternating before/after rounds produced 84 fresh processes. All YAML hashes within each scenario match.

The minimal configuration uses Rule mode with MATCH,DIRECT, DNS/TUN disabled, and a loopback controller only for readiness, rule-count and drain verification. TCP scenarios also enable loopback SOCKS. A 15-second warmup precedes a three-second sample window at approximately 250 ms intervals. Each observation is a per-process phase median; the table takes the median of seven processes, not hundreds of falsely independent polling samples.

All values below are MiB. Working-set parentheses show the full range across seven processes. Private bytes are Windows private committed memory, a different metric from working set.

ScenarioBefore working set (range)After working set (range)Private bytes: before → after
Idle19.19 (18.97–19.29)19.17 (19.07–19.27)50.32 → 50.43
100k noncoalescing CIDR /32 rules26.26 (26.04–29.70)25.77 (25.47–29.66)56.31 → 55.71
100k synthetic MRS domains30.63 (30.48–30.74)30.65 (30.55–30.83)60.94 → 60.96
100k synthetic GeoSite domains92.38 (57.52–93.68)70.30 (45.58–76.59)122.50 → 100.18
100 held TCP connections29.71 (29.50–29.86)29.73 (29.01–29.89)60.11 → 60.18
100 TCP connections, 60 seconds after close30.55 (30.46–30.63)30.58 (30.25–30.71)61.00 → 61.17
1,000 held TCP connections113.43 (110.61–115.91)115.48 (113.11–116.14)144.85 → 147.43
1,000 TCP connections, 60 seconds after close122.22 (121.41–122.99)122.69 (121.96–122.95)157.06 → 157.17
  • GeoSite median working set fell 23.9%, private bytes 18.2%. Ranges still overlap; this does not guarantee a fixed saving on every startup.
  • The independent loopback echo workload verifies every SOCKS handshake and echo, then transfers 128 × 4096 bytes in each direction per connection, with at most 16 workers. At 1,000 connections that is 500 MiB each way. Every connection closes and the controller reports zero live connections. Working set during traffic was 117.02 → 117.34 MiB; after traffic while still held, 122.25 → 122.71 MiB; 15 seconds after close, 122.22 → 122.69 MiB.
  • TCP process RAM was not shown to improve or return to idle within 60 seconds. The 1,000-held-connection median was actually +1.8%. Releasing objects or returning buffers to pools, and lower B/op, do not mean immediate return of memory to the OS.
  • Short-term CIDR working set was unstable: an earlier independent seven-round cohort at 4a5d6937 was 26.27 → 29.44 MiB (+12.1%); the final cohort was −1.9%. The intervening code changes only touched binary export, which this scenario does not execute, and ShadowQUIC lint. Do not attribute that change to a process-RAM fix. Separate, explicitly forced-GC heap diagnostics confirmed that approximately 4.58 MiB of duplicate IPRange storage was no longer live. This verifies unreachability, not natural process-memory savings.

Identical-source CPU/allocation A/B ​

The main matrix compares 4a59a634 → 4a5d6937: 14 packages, 38 cases, seven rounds, 532 samples. Each concrete case/sample runs in a fresh process with identical benchmark-only source, GOMAXPROCS=1, -test.cpu=1, and a target of at least two seconds. After ca548d9f fixed CIDR export, four CIDR cases were rerun against final 399e7f83, adding 56 samples. The other 34 cases were not rerun against the final hash; their results remain explicitly attributed to 4a5d6937.

WorkBefore → afterAllocated bytes / allocations: before → after
DNS IPv4 cache hit675.8 → 178.7 ns (−73.6%)512 / 11 → 88 / 3
Deadline refresh171.70 → 61.61 ns (−64.1%)128 / 2 → 0 / 0
TCP tracker with traffic control2468.0 → 850.5 ns (−65.5%)7680 / 18 → 688 / 9
Traffic-control session: global1372.0 → 122.5 ns (−91.1%)7200 / 11 → 208 / 2
Single-destination UDP mapping creation931.7 → 547.6 ns (−41.2%)3796 / 11 → 2788 / 7
UDP destination lookup24.89 → 16.99 ns (−31.7%)0 / 0 → 0 / 0
SOCKS UDP handleSocksUDP286.6 → 221.5 ns (−22.7%)596 / 6 → 208 / 3
Domain-suffix constructor64.25 → 32.40 ns (−49.6%)64 / 2 → 32 / 1
Kernel-direct address eviction76.81 → 60.93 µs (−20.7%)47526 / 27 → 47156 / 25
CIDR export, 10k (final-candidate rerun)230.28 → 53.92 µs (−76.6%)655488 / 6 → 655488 / 6

These timing improvements have Mann–Whitney U p=0.001, except SOCKS UDP at p=0.011; no multiple-comparison correction was applied. Full ranges and every raw sample are in the evidence archive. This is a development desktop, not an isolated, fixed-frequency server.

Costs retained and improvements not claimed ​

  • Default TCP tracker: 352 → 256 B/op, still 3 allocations; 308.5 → 295.5 ns (p=0.620). UnwrapReader instead rose 3.337 → 4.701 ns, +40.9% (p=0.001), still allocation-free.
  • A second UDP destination had a +20.2% timing median (p=0.259, not significant); two or more destinations cost another 144 B. Timing at 100/1,000 mappings was unchanged. The single-destination win is not the whole story.
  • VLESS 0-RTT: 17616 → 16434 B/op, but 27 → 29 allocations; 89.59 → 89.85 µs (p=1.000). No handshake-speed claim.
  • NAT hit +2.4% (p=0.026); DoT median +5.0% (p=0.383); ReadCached median +2.1% (p=1.000), saving one allocation only. TCP relay, sing handlePacket, and sequential xHTTP showed no significant speedup.
  • Original CIDR export was 226.6 → 387.0 µs (+70.8%). That regression remains recorded. The correction encodes compact tables directly and passes independent wire-reference, mixed-family, mapped-address and boundary tests.
  • Different new sequential GeoSite attribute variants re-decode the raw list; the same-matcher cache remains. Deterministic fake-loader tests show two loads for a pair instead of one. Full-size sequential-variant CPU A/B was not measured: this trades load-time CPU for lower retained storage.
  • Queue/cache limits, TTLs and features are not reduced; no GC tuning. Kernel-DIRECT TC eBPF remains disabled. Real-device TUN/OpenWrt RAM and WAN were not retested. sync.Pool is not a permanent-leak fix or a capacity bound guaranteed by GOMAXPROCS.

Evidence and validation ​

Download raw results, identical-source harnesses and reproduction notes. Includes both process cohorts, both microbenchmark revisions, pre-fix regressions, fixture generators and binary/config/harness hashes; no executables or private agent conversations.

Final-candidate go test -p=1 -count=1 ./... passed. Scoped race tests, vet, VMess with_low_memory and gofmt passed; golangci-lint 2.12.2 reported 0 issues with CGO both disabled and enabled. The actual validation toolchain was Go 1.26.3, not a claimed Go 1.20 test run. These local checks do not prove remote CI or deployment status.

2026-08-28 core hot-path optimization A/B ​

This run compares Aster 90f0e4ee (before this wave) with 72048c8a (performance commit 9824ccdb plus its lint fix). Each revision was built as separate Windows amd64 test binaries. Every sample used a fresh process, GOMAXPROCS=1, -test.cpu=1, GOAMD64=v1, Go 1.26.3, and a two-second run. Before and after were interleaved for seven rounds, alternating which revision ran first.

The main matrix covered 11 packages and 24 identically named cases per revision, for 168 samples per revision. The table reports the seven-round median and full range. Every improvement listed below had p=0.001 in benchstat's Mann–Whitney U test. The machine was Windows 11, a Ryzen 9 5900X (12C/24T), and 64 GB DDR4. Total CPU samples were 9.85–17.22% before the run and 7.19–15.71% after it. This is still a development-machine microbenchmark, not WAN throughput.

Aster core workBefore median (range)After median (range)ResultAllocations
Existing Kernel DIRECT flow refresh221.0 ns (208.2–243.9)182.0 ns (179.8–185.9)1.21×; time down 17.6%0 B/0 → 0 B/0
Existing NAT flow lookup64.26 ns (63.99–65.38)11.67 ns (11.58–12.09)5.51×; time down 81.8%0 B/0 → 0 B/0
UDP WriteBack target update7.515 ns (7.491–7.802)3.918 ns (3.908–4.084)1.92×; time down 47.9%0 B/0 → 0 B/0
DomainSet.Has (short)166.7 ns (165.0–170.5)56.32 ns (54.96–56.96)2.96×; time down 66.2%0 B/0 → 0 B/0
Miss in 100k merged CIDRs84.63 ns (83.60–90.87)12.93 ns (12.57–15.84)6.55×; time down 84.7%0 B/0 → 0 B/0
GetUser (10,000 users)56.67 ns (55.23–61.04)40.26 ns (37.26–47.98)1.41×; time down 29.0%0 B/0 → 0 B/0
Upload traffic increment3.561 ns (3.533–3.652)1.742 ns (1.717–1.852)2.04×; time down 51.1%0 B/0 → 0 B/0
TCP tracker lifecycle574.7 ns (561.3–592.6)257.6 ns (253.5–314.1)2.23×; time down 55.2%528 B/8 → 352 B/3
Default-rule match22.97 ns (22.78–23.30)20.13 ns (20.03–20.53)1.14×; time down 12.4%0 B/0 → 0 B/0
UDP handlePacket (same benchmark-only harness)299.8 ns (259.1–353.4)187.4 ns (170.5–230.0)1.60×; time down 37.5%280 B/8 → 208 B/2

Controls and results that are not advertised as wins ​

  • The production UDP metadata-pool code is identical in both revisions. The ordinary test binaries initially showed 11.50 → 13.18 ns. With one minimal, identical harness on both revisions it measured 12.84 → 12.73 ns (p=0.874), confirming a test-layout artifact rather than a pool regression. The complete handlePacket path did remove 37.5% of the time and 75% of allocation count.
  • Inserting 100 or 1,000 UDP mappings did not change significantly (1,000: 348.2 → 331.1 µs, p=0.097); allocation stayed at 615,856 B/2,049 allocs.
  • The 1 KiB, 16 KiB, and 64 KiB AnyTLS WriteDataFrame cases all stayed zero-allocation. The 1 KiB and 16 KiB medians were flat; the 64 KiB change was not significant (p=0.073).
  • The shared 32 KiB relay median moved from 3.766 to 3.959 µs (+5.1%), and the tunnel TCP relay from 3.899 to 4.023 µs (+3.2%). Both before/after ranges overlap. This run does not claim a desktop TCP speedup and does not overwrite the OpenWrt result below. Determining whether 3–5% is a real regression requires a fixed-frequency CPU or another OpenWrt hardware rerun.

This A/B validates Aster's internal hot paths in 9824ccdb. It did not run a Mihomo binary or use an external endpoint, so it answers “did this Aster commit lower core overhead,” not “how much faster will a download be.”

2026-08-24 review-wave validation ​

The unreleased working tree based on 8462a265 was rebuilt against Mihomo v1.19.30 (ac017cdd) as separate Windows amd64 test binaries. Every sample used a fresh process, GOMAXPROCS=1, -test.cpu=1, GOAMD64=v1, Go 1.26.3, and a two-second run; Aster and Mihomo were interleaved for seven rounds. This is regression validation on the 5900X development machine. It does not replace the OpenWrt primary result below and is not WAN throughput.

Core workMihomo 1.19.30 median (range)Aster median (range)Allocations
UDP metadata54.44 ns (49.80–71.77)12.22 ns (11.76–12.48), 4.45×416 B/1 → 0 B/0
Disabled debug log206.3 ns (204.3–245.5)2.132 ns (2.096–3.063), 96.8×24 B/1 → 0 B/0
AnyTLS frame, 1 KiB72.55 ns (66.80–82.65)31.67 ns (30.93–33.76), 2.29×64 B/1 → 0 B/0
AnyTLS frame, 16 KiB322.8 ns (320.4–354.0)244.3 ns (218.4–277.3), 1.32×64 B/1 → 0 B/0
Shared 32 KiB relay4.038 µs (4.005–4.157)3.916 µs (3.763–4.477)64 B/1 → 0 B/0
Tunnel 32 KiB TCP relay4.079 µs (3.892–4.666)4.016 µs (3.840–4.198)64 B/1 → 0 B/0

The two TCP relay ranges overlap, so this run establishes that Aster stayed zero-allocation, not a statistically significant desktop TCP speedup. UDP, logging, and AnyTLS allocation classes and direction agree with the OpenWrt result.

Same-host before/after regression benchmarks for this review wave showed:

Aster-only hot pathBefore reviewAfter reviewResult
Existing Kernel DIRECT flow refresh15.286 µs; 64 B/1 alloc270.8 ns; 0 B/0 alloc56.4×; no full scan before the next expiry
Miss in 100k merged, disjoint CIDRs2.385 ms; 0 alloc114.7 ns; 0 allocabout 20,790×; restored binary search
Insert 1,000 UDP-association mappings41.31 ms; 46.37 MB486.3 µs; 542.6 KiB84.9×; removed whole-map copying

For five full-core runs with the same minimal config, sampled 15 seconds after startup, Windows working-set medians were 19.00 MiB for Aster and 18.07 MiB for Mihomo (Aster +0.93 MiB / +5.2%). Private bytes were 52.52 MiB versus 51.29 MiB (+1.23 MiB / +2.4%). These Windows counters cannot be mixed with Linux RSS/PSS and again do not support a general claim that Aster uses less RAM when idle.

OpenWrt on real hardware (primary result) ​

The comparison baseline is Mihomo v1.19.30, commit ac017cdd246ce8bd547653d927e7bf77d7ee73d5. Aster was main at 0590d3a4 (this review/fix wave plus the 128 KiB frame pool). Each side was built separately with the same Go version, target, flags, and benchmark harness as a Linux amd64 go test -c binary, then run for three sequential rounds, at least 2 seconds each. The table uses the three-round median. Rerun on 2026-08-19 on the same OpenWrt soft router after the fixes landed.

EnvironmentActual value
SystemOpenWrt, Linux 6.6.86, x86-64
CPURyzen 7 5825U host, VM configured with 12 vCPU
CPU frequencyGuest sample averaged 2.342 GHz before the unrestricted run. This is /proc/cpuinfo; the hypervisor does not guarantee a constant clock
MemoryVM configured with about 6 GB (MemTotal 6081752 kB)
DRAM frequencyHypervisor did not expose it, unavailable
Go1.26.3, GOAMD64=v1 cross-compiled Linux amd64 test binary
Load average before the test0.07/0.03/0.00 (1/5/15 minutes)
Load average after the test0.92/0.37/0.14 (1/5/15 minutes)
Core workMihomo 1.19.30 medianAster medianAster relative result
UDP packet metadata70.00 ns; 416 B/1 alloc12.89 ns; 0 B/0 alloc5.43× faster; removed the 416 B allocation
Disabled debug log228.0 ns; 24 B/1 alloc2.254 ns; 0 B/0 alloc101× faster; removed the event allocation
AnyTLS frame (1 KiB)71.54 ns; 64 B/1 alloc35.98 ns; 0 B/0 alloc1.99× faster; latency down 50%
AnyTLS frame (16 KiB)267.1 ns; 64 B/1 alloc196.8 ns; 0 B/0 alloc1.36× faster; latency down 26%
AnyTLS frame (64 KiB)no matching bench1.148 µs; 0 B/0 allocThe isolated WriteDataFrame/65536 fixture is zero-alloc with the 128 KiB pool
TCP relay (32 KiB)4.546 µs; 7.21 GB/s; 64 B/1 alloc4.449 µs; 7.37 GB/s; 0 B/0 allocLatency down 2.1%; throughput up 2.2%

Three-round ranges ​

BenchmarkMihomo 1.19.30 rangeAster range
UDP packet metadata69.74–72.12 ns/op12.84–13.00 ns/op
Disabled debug log223.6–231.5 ns/op2.230–2.270 ns/op
AnyTLS frame (1 KiB)71.35–71.69 ns/op35.76–36.89 ns/op
AnyTLS frame (16 KiB)264.9–271.8 ns/op195.2–216.4 ns/op
TCP relay (32 KiB)4.538–4.576 µs/op4.423–4.505 µs/op

This 5825U soft router is still stronger than many MT7621, low-end ARM, or cheap VPS hosts. These results only prove the optimizations still work on real OpenWrt. Weaker devices will have lower absolute GB/s, and the improvement ratio must be remeasured on that device.

Hyper-V low-resource simulation (secondary result) ​

To get closer to weaker hardware most users have, we added another resource limit on the same Hyper-V OpenWrt VM. The benchmark was allowed 1 vCPU, and at most 25 ms of execution every 100 ms, which is single-core 25% CPU time. Memory was capped at 512 MiB, with swap disabled. Rerun on 2026-08-19 against Mihomo v1.19.30 with the same Linux amd64 test binaries.

This simulates a resource-starved situation. It is not an MT7621 or ARM ISA simulation. It answers “do Aster’s optimizations still exist when CPU is slow and RAM is scarce,” but it does not replace a test on a real low-end router.

EnvironmentActual value
SystemOpenWrt on Hyper-V, Linux 6.6.86, x86-64
Host CPURyzen 7 5825U
CPU limitPinned to vCPU 0; cgroup cpu.max = 25000 100000, i.e. single-core 25% CPU time
Guest-reported CPU frequencyAverage 2.342 GHz before the unrestricted segment. This is a frequency sample and does not include the 25% CPU-time limit
Memory limit512 MiB, no swap
DRAM frequencyHypervisor did not expose it, unavailable
Go1.26.3, Linux amd64 test binary
Load average before the test0.07/0.03/0.00 (same start as the unrestricted segment)
Load average after the test0.31/0.31/0.14 (1/5/15 minutes)
Throttle confirmation3,660 of 3,709 CPU quota periods were throttled

Both versions ran the same benchmarks sequentially, three rounds, at least 2 seconds each. The table uses the three-round median:

Core workMihomo 1.19.30 medianAster medianAster relative result
UDP packet metadata251.5 ns; 416 B/1 alloc48.01 ns; 0 B/0 alloc5.24× faster; removed the 416 B allocation
Disabled debug log854.3 ns; 24 B/1 alloc8.054 ns; 0 B/0 allocabout 106× faster; removed the event allocation
AnyTLS frame (1 KiB)331.0 ns; 64 B/1 alloc139.6 ns; 0 B/0 alloc2.37× faster; latency down 58%
AnyTLS frame (16 KiB)1.080 µs; 64 B/1 alloc790.8 ns; 0 B/0 alloc1.37× faster; latency down 27%
TCP relay (32 KiB)17.644 µs; 1.86 GB/s; 64 B/1 alloc16.812 µs; 1.95 GB/s; 0 B/0 allocLatency down 4.7%; throughput up 5.0%

Peak memory ​

Keep two “memory” numbers separate:

SituationMihomo 1.19.30AsterConclusion
Full core, minimal profile, idle 15 s RSS34.7 MiB39.3 MiBAster is 4.6 MiB larger, about +13%

The full-core test used the same profile: Direct mode, no proxies, no rules, silent logs, IPv6 off, TUN off, bound only to unused 127.0.0.1 ports. Aster and Mihomo each ran three rounds, waiting 15 seconds after each start. RSS three-round ranges were Aster 38.5–39.3 MiB and Mihomo 34.7–35.0 MiB. The OpenWrt kernel did not provide smaps_rollup, so there is no PSS. Only /proc/<pid>/status VmRSS is reported.

So you cannot say Aster uses less RAM when idle. Aster ships more features and code, and the minimal idle base is currently about 4.6 MiB larger than Mihomo 1.19.30. The isolated-process peaks below are Max RSS of a Go test program that only loaded that package. They are not the total RAM of a full proxy with rules, GeoIP, DNS cache, and live connections.

Isolated-process peaks for core hot paths ​

Each case was rerun as an isolated process for three rounds. Max RSS was recorded with Linux /usr/bin/time -v. The table uses the three-round median. Both sides used the same single-core 25% CPU time and 512 MiB memory limit.

Core workMihomo 1.19.30 peak RSSAster peak RSSDifference
TCP relay (32 KiB)40.5 MiB51.8 MiBAster is 28% larger
Disabled debug log33.5 MiB23.0 MiB31% less
UDP packet metadata53.0 MiB53.4 MiBEven
AnyTLS frame (1 KiB)35.4 MiB26.0 MiB27% less
AnyTLS frame (16 KiB)34.8 MiB25.5 MiB27% less

After the 1.19.30 rerun, do not keep saying all five cases used 30–49% less RAM. Disabled log and AnyTLS are still clearly lower. UDP is even. The isolated TCP test process peaked higher for Aster. This compares temporary pressure from that package test, not Aster’s RAM under every configuration. The Aster test binary is larger than Mihomo’s, so RSS includes the whole Go test process.

Three-round ranges: TCP Mihomo 39.5–40.5 MiB, Aster 51.5–52.0 MiB; log Mihomo 32.5–33.5 MiB, Aster fixed 23.0 MiB; UDP Mihomo 50.8–53.9 MiB, Aster 53.0–54.9 MiB; AnyTLS 1 KiB Mihomo 35.0–35.5 MiB, Aster fixed 26.0 MiB; AnyTLS 16 KiB Mihomo 34.7–35.1 MiB, Aster 25.5–26.0 MiB. All tests used 0 swap.

Constrained three-round ranges ​

BenchmarkMihomo 1.19.30 rangeAster range
UDP packet metadata244.3–256.4 ns/op47.85–48.21 ns/op
Disabled debug log854.1–856.9 ns/op7.984–8.157 ns/op
AnyTLS frame (1 KiB)323.4–335.5 ns/op135.5–155.2 ns/op
AnyTLS frame (16 KiB)1.075–1.092 µs/op772.3–801.6 ns/op
TCP relay (32 KiB)17.598–17.756 µs/op16.771–16.841 µs/op

Absolute processing time became about four times the unrestricted run, which shows the CPU quota was real. Aster was still faster on all five latency jobs. In this CPU-quota experiment on the same x86 VM, TCP improved from 2.1% unrestricted to 4.7%, and AnyTLS 1 KiB from 1.99× to 2.37×. The UDP ratio eased from 5.43× to 5.24×. This cannot be generalized into “every weaker machine improves more”; ARM/MIPS, cache, and memory-bandwidth effects still require native testing. The homepage still uses the unrestricted OpenWrt TCP about 2% versus Mihomo 1.19.30, not this more flattering 4.7%, as the main advertised number.

Per-protocol loopback tests ​

A faster core microbenchmark does not mean every protocol’s end-to-end throughput rises in lockstep. To check that, we previously connected Aster and Mihomo to the same local server on a Linux Docker host network and measured read/write with 16 KiB chunks. The table below is still versus Mihomo v1.19.29; the protocol servers were not rerun against v1.19.30.

ProtocolServerAster vs Mihomo
Direct (control)noneRead/write gap within about 3%, treated as even
Shadowsocks AES-256-GCMshadowsocks-rustSame allocation count; write samples were too noisy, no speedup claim
VMessV2Fly 4.45.2Write about 645 vs 647 MB/s; read about 488 vs 488 MB/s, even
VLESS + TLSV2Fly 4.45.2Typical interleaved three-round gap about -2% to +3%, even
Trojan + TLStrojan 1.16Interleaved five-round median: write 335 vs 341 MB/s, read 313 vs 314 MB/s, even
Snell v3 + HTTP obfsOpenSnell 3.0.1Interleaved three-round median: write 429 vs 430 MB/s, read 763 vs 774 MB/s, even
Hysteria v1Hysteria 1.3.5, 100 MbpsBoth about 12.25–12.28 MB/s, limited by the server bandwidth cap
AnyTLSNo same-condition loopback server todayDo not mix in an external node; only the earlier frame microbenchmark is listed

Numbers in the table are Aster vs Mihomo. VMess, VLESS, Trojan, Snell, and Shadowsocks B/op and allocs/op are also roughly the same. This round of work mainly improved shared relay, UDP metadata, logging, and AnyTLS frame hot paths. It did not magically make every encrypted protocol much faster.

These protocol tests are still the earlier Docker loopback against Mihomo v1.19.29. They were not rerun against v1.19.30. This 5900X host has Docker Engine, but not the original pinned protocol server images, and pulling new tags would not match those SHAs. The encryption and copy paths are shared; the 1.19.30 microbenchmarks already cover the changed relay, UDP, log, and AnyTLS frame work. The table is only here to show that encrypted protocols did not all jump together. They are not homepage advertising numbers.

Server images were pinned to shadowsocks-rust sha256:85d01d…e1359, V2Fly sha256:e81a07…de78c, Trojan sha256:5b36c2…b98b5f7, OpenSnell sha256:70053f…345467, and Hysteria v1.3.5 sha256:4c8c92…f1e35.

UDP metadata allocation before and after ​

The same benchmark measures Mihomo’s direct metadata-construction path and Aster’s object-pool path, so they can be compared on the same machine:

PathTimeMemory allocation
Mihomo 1.19.30 constructs per packet69.74–72.12 ns/op416 B/op, 1 alloc/op
Aster metadata pool12.84–13.00 ns/op0 B/op, 0 allocs/op

The object-pool path is about 5.43× faster on the three-round median and removes a 416-byte heap allocation per packet.

High-end development machine (appendix) ​

The original Aster absolute numbers were taken on a high-end desktop. They are only useful as a development-time regression check and should not represent an ordinary router. On 2026-08-19 the same 5900X reran the same microbenchmarks against Mihomo v1.19.30 and Aster main, interleaved, three rounds, at least 2 seconds each:

EnvironmentActual value
SystemWindows 11 amd64; balanced power mode
CPURyzen 9 5900X, 12 cores/24 logical CPUs
Reported frequencyWindows \Processor Frequency stayed at 3701 MHz (advertised clock, not actual boost)
DRAM64 GB (4×16 GB) G.Skill DDR4-3600, configured 3600 MT/s
Go1.26.3 windows/amd64
CPU load before each caseAbout 30% average across start samples (17.6–60.9%)
CPU load after each caseAbout 34% average across end samples (20.3–54.8%)
Processor queue0 throughout

Background load is still not low, so the 5900X run is not the homepage’s primary comparison. It is only here to confirm the desktop development machine did not regress. An earlier same-day pass with about 42–47% background load and no interleaving produced contradictory TCP numbers and was discarded. After the fixes landed, an interleaved rerun showed:

Core workMihomo 1.19.30 medianAster medianAster relative result
UDP packet metadata152.4 ns; 416 B/1 alloc11.91 ns; 0 B/0 alloc12.8× faster; removed the 416 B allocation
Disabled debug log455.3 ns; 24 B/1 alloc2.268 ns; 0 B/0 alloc201× faster; removed the event allocation
AnyTLS frame (1 KiB)74.99 ns; 64 B/1 alloc34.70 ns; 0 B/0 alloc2.16× faster
AnyTLS frame (16 KiB)260.0 ns; 64 B/1 alloc184.0 ns; 0 B/0 alloc1.41× faster
AnyTLS frame (64 KiB)no matching bench1.056 µs; 0 B/0 allocThe isolated WriteDataFrame/65536 fixture is zero-alloc
TCP relay (32 KiB)10.044 µs; 3.26 GB/s; 64 B/1 alloc11.436 µs; 2.87 GB/s; 0 B/0 allocRanges overlap; Aster’s median was slower this pass, no 5900X TCP speedup claim

Three-round ranges: UDP Mihomo 145.8–155.1 ns, Aster pool 11.69–12.21 ns; log Mihomo 442.9–456.5 ns, Aster 2.227–2.352 ns; AnyTLS 1 KiB Mihomo 74.12–77.73 ns, Aster 33.21–36.36 ns; AnyTLS 16 KiB Mihomo 257.9–271.6 ns, Aster 182.9–200.4 ns; TCP Mihomo 9.963–11.705 µs, Aster 10.394–11.947 µs. The 32 KiB Relay32KiBComparison median was Aster 9.332 µs versus Mihomo 10.775 µs. Desktop TCP is noisy under background load; the homepage still uses OpenWrt.

How to rerun ​

Run the current Aster hot-path suite from the repository root:

sh
GOAMD64=v1 GOMAXPROCS=1 go test \
  ./component/kerneldirect ./component/nat ./component/trie ./component/cidr \
  ./component/aster ./listener/sing ./tunnel ./tunnel/statistic \
  -run '^$' \
  -bench 'Benchmark(ObserveFlowRefresh|WriteBackProxyUpdate|TableExistingFlow|DomainSetHas|IpCidrSetMergedMiss|ManagerGetUser|ManagerPushUploaded|TCPTrackerLifecycle|MatchDefaultRule|PacketMetadata)$' \
  -benchmem \
  -benchtime=2s \
  -count=7 \
  -cpu=1

For a commit-to-commit A/B, build each revision with go test -c, launch a fresh test process for every sample, and alternate whether before or after runs first. Do not mix compilation into a still-rebuilding go test timing, and do not compare one-off samples from different machines.

The older Aster-versus-Mihomo suite is:

sh
go test \
  ./common/net ./component/nat ./constant ./listener/sing ./log \
  ./transport/anytls/padding ./transport/anytls/session \
  ./tunnel ./tunnel/statistic \
  -run '^$' \
  -bench 'Benchmark' \
  -benchmem \
  -count=3

When comparing against Mihomo, build the same Linux amd64 test binaries from tag v1.19.30 (ac017cdd) and run three sequential rounds on the same machine. Do not compute a percentage from a single run on different machines.

How to read the numbers ​

  • These are in-process microbenchmarks. They mainly measure Aster Core’s own extra cost.
  • TCP and AnyTLS GB/s numbers are memory / net.Pipe paths, not real WAN throughput.
  • Real proxy speed is still limited by encryption, RTT, loss, MTU, NIC, OS, CPU architecture, and the server.
  • The isolated WriteDataFrame fixture is zero-allocation for its 1 KiB, 16 KiB, and 64 KiB cases (64 KiB uses the 128 KiB pool). This does not mean every complete AnyTLS-session path is zero-allocation. Larger or misaligned buffers may still allocate.
  • Microbenchmarks are best at catching regressions. They do not promise the same network speed on every device.

A real-hardware counter-example for TC eBPF ​

Kernel DIRECT itself does not require TC eBPF. The recommended OpenWrt path is the nftables learned exclude set handing DIRECT back to Linux forwarding/NAT, while keeping flow offload. Experimental kernel-direct-ebpf looks up generation, IPv4/IPv6 LPM, a 40-byte 5-tuple LRU, and per-CPU counters on every LAN ingress packet.

On the same router and the same Speedtest server ID 37639 A/B:

StateDownload
TC eBPF on692,335,768 bps
TC filters temporarily unloaded1,647,299,448 bps
TC persistently off, after restart1,643,651,288 bps

After unload it is about 2.37× the enabled result, and it returns to that network’s original ~1.7 Gbps class. The main reason is that this TC ingress hook interferes with OpenWrt flow offload, plus the classifier itself runs per packet. It does not mean every other eBPF program or every other piece of hardware is slower.

This is also why microbenchmarks and real network results must be kept separate: a fast map lookup or data structure does not mean putting it on every ingress packet, and changing the kernel offload path, will raise end-to-end throughput. Before and after enabling TC, keep the client, server ID, protocol, and time window fixed and run multiple times. If there is no clear gain, leave kernel-direct-ebpf: false.