nak / mitra vibe fuzz

Last active 3 hours ago

Like 0

Findings from a vibe fuzzing session I can PR some fix attempts here too if you want, but wanted to at least pass the remaining 2 findings upstream

Revision d512e535ec0fed7ccae830da1e69dcce3973d09c

fuzz_findings.md Raw

Remaining fuzzing findings for Mitra

Session date: 2026-08-02. Tested at 8b5c382b (Mitra v5.8.0) with cargo-fuzz, libFuzzer, and AddressSanitizer.

The previously reported clean_title byte/character mismatch is intentionally not included here. Its fix is being handled in PR #254. The two observations below remain; neither produced an escaping bypass or memory-safety issue in the tests performed.

# Finding Severity Introduced First released
1 clean_title truncation splits HTML entities Low 7b2b46d5 v2.14.0
2 clean_remote_content is not idempotent after one pass Informational dependency behaviour n/a

1. clean_title truncation splits HTML entities

Impact. A truncated title can contain a partial HTML entity and render as corrupted text. Cleaning is not idempotent. No escaping bypass was observed; the fragment remains escaped.

Mechanism. clean_title first calls clean_html_all, which escapes < as &lt;, and then truncates the resulting string by characters. The truncation can land inside an entity, leaving a suffix such as &l. A second cleaning pass escapes that bare & as &amp;, so the result changes again.

Reproducer:

use mitra_validators::posts::clean_title;

let input = format!("aaa{}", "<".repeat(75));
let once = clean_title(&input);
let twice = clean_title(&once);

assert_ne!(once, twice);

Observed tails:

pass 1: ...lt;&lt;&l...
pass 2: ...lt;&lt;&a...

History. Commit 7b2b46d5 (2024-03-23, v2.14.0), "Remove HTML tags from post title", changed the first step from title.trim() to clean_html_all(title).trim().to_owned(). This placed escaped entities in the string that was already being truncated by character count.

Possible handling. Ensure truncation cannot cut an &...; sequence, or truncate before HTML escaping in a way that preserves the sanitizer's output guarantees. The character-count change from PR #254 does not address this separate behavior.

2. clean_remote_content is not idempotent after one pass

Impact. Informational. The observed input converges by the third pass and no dangerous markup survives. This is recorded because sanitizer output that re-parses into a different tree has the same general shape as mutation-XSS issues, even though no exploit was found here.

Mechanism. An unknown tag leaves the HTML parser in a state where a following heading is re-nested on the next pass.

Input:

v555<h4><hh4><h3>55\t

Output over repeated calls to clean_remote_content:

pass 1: v555<h4><h3>55\t</h3></h4>
pass 2: v555<h4></h4><h3>55\t</h3>
pass 3: v555<h4></h4><h3>55\t</h3>

clean_remote_content uses clean_html, which relies on ammonia's default tag allowlist. The h3 and h4 tags are not from Mitra's CONTENT_ALLOWED_TAGS; that list applies only to the strict local-content path.

History. This was not traced to a Mitra defect. It is html5ever's reparse behavior for invalid heading nesting, exposed through ammonia. The tested versions were ammonia 4.1.4 and html5ever 0.39.0. clean_remote_content itself dates to d7750442 (2025-02-26, v3.18.0), "Refactor content cleaners", while the underlying parser behavior predates it.

Current fuzz-harness handling. The html_sanitize target asserts convergence within eight passes and checks that no executable tags survive, rather than requiring one-pass idempotence. This keeps the target useful while retaining coverage of the known behavior.