Project homepage Mailing List  Warmcat.com  API Docs  Github Mirror 
    npro  
 Modern all-safe Rust Network Protocol library supporting h1, h2, h3, ws, wt sans-IO and with socket IO + tls
git clone https://npro.rs/repo/npro
 
root / fuzz / seeds / chunked / unfinished.txt
Author[]Andy Green <andy@warmcat.com> 2026-10-04 20:20 UTC
Committer[]Andy Green <andy@warmcat.com> 2026-10-05 06:32 UTC
Tree95beb7f39748466f9b37c772a79a4b558f6602fa   Raw Patch
 
npro-h1: the h1 head parser, its header table and C's token set
npro-h1: the h1 head parser, its header table and C's token set

Phase 1c's crate, no_std with no dependencies: C's
lib/sansio/http/parsers.c, the part of it that parses heads.

 - token: C's enum lws_token_indexes, its 97 tokens at C's indices with
   C's spellings, with every header option on as C's default build has
   them.  C matches names with a generated trie, the lextable; here it is
   lookup() over the spellings.  No spelling is the start of another, a
   test says so, so a name matches exactly when the trie would reach its
   terminal, and what the trie is (an encoding) is not ported.

 - table: C's ah, HeaderTable<S> over caller-owned storage, [u8; 4096] on
   a device or a Box<[u8]>, at most 32768 bytes as C caps it.  It is laid
   out as C's is, so that it fills at the same byte: each value with its
   NUL, each unknown header as an eight byte record and its name and
   value, the byte C leaves unused after a path's NUL, and C's 97 fragment
   slots.  Where C says present by a nonzero fragment index, here it is
   Slot::Present.  A client's own request goes in with create(), C's
   lws_hdr_simple_create(), but all or nothing; snapshot() and rewind()
   drop an interim response, as C's rx snapshot and rewind.

 - head: C's lws_parse() and lws_parse_urldecode(), branch for branch, a
   byte at a time and restartable at any byte.  C's parser_state is State,
   the URI decoder's ues, ups and post_literal_equal are Esc, Path and
   Arg, and the name matcher's lextable_pos and unk_pos are Name.  Each of
   C's ways out (bad_request_line, forbid, too_large, bare_lf...) is a
   Cause, and Refused says what a server answers for it, C's LPR_REFUSED
   status, or nothing for LPR_FAIL.  Config holds C's token_limits and
   what a server does with an unknown method (refuse, or C's fallback
   role).  What C does on purpose stays, the version on a request line
   held to its target's limit among it.

 - fields: C's Content-Length reader, digits with trailing spaces and no
   wrapping, and its test for a Transfer-Encoding that is exactly one
   chunked coding.

Porting it found three bugs in C, fixed there first (lws _temp
f5338775b, 847081657 and 3c8459075), so this follows C as it now is:

 - C knew a name had begun by its record's offset, unk_pos, being
   nonzero, but a server's head starts at offset 0.  So on a request's
   first token C began the name again at the second byte, losing nine
   bytes of the ah, and took a LF as the second byte as a bare LF,
   closing without the 400 a LF third gets.  Both C and npro now know a
   name's start by the matcher being at its start: here, a state of its
   own.

 - A strict server took a header line starting with a bare CR, then not
   LF, into an unknown header's name, "\rx-a:": the CR matched the start
   of the empty line's "\r\n", and only the byte the name stopped
   matching at was checked.  Now every byte of the name before it is
   checked too, and the CR refuses the request: what is in front of us
   may have ended the line there.

 - A repeated header's value kept its own leading spaces, "a,   b",
   since the swallow looked at the header's first piece, not the one
   being filled, and knew only SP.  RFC 9110 5.5's OWS around a value is
   not part of it: a value's leading SP and HT are dropped while its
   piece holds nothing of its own, a repeat's only the SP joining it to
   the one before.  The request line keeps its rule, extra SPs before the
   target and the version dropped, and a HT there no SP.

The crate is added to deny.toml's allowlist, and to the no_std builds of
ci.sh and sai, where it builds for cortex-m0, cortex-m4f and riscv32imc.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019kg5Eemy68ZaqDBcUJQG6J
diff --git a/Cargo.lock b/Cargo.lock index 44fd6df..058bc2d 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -19,6 +19,10 @@ dependencies = [ ] [[package]] +name = "npro-h1" +version = "0.0.2" + +[[package]] name = "npro-test" version = "0.0.2" dependencies = [ diff --git a/crates/npro-h1/Cargo.toml b/crates/npro-h1/Cargo.toml new file mode 100644 index 0000000..d5b32c8 --- /dev/null +++ b/crates/npro-h1/Cargo.toml @@ -0,0 +1,15 @@ +[package] +name = "npro-h1" +description = "The sans-IO h1 of npro: the head parser, chunked coding, and in time the h1 client and server" +readme = "../../README.md" +keywords = ["http", "sans-io", "no-std", "parser"] +categories = ["network-programming", "parser-implementations", "no-std"] +version.workspace = true +edition.workspace = true +rust-version.workspace = true +license.workspace = true +homepage.workspace = true +repository.workspace = true + +[lints] +workspace = true diff --git a/crates/npro-h1/src/fields.rs b/crates/npro-h1/src/fields.rs new file mode 100644 index 0000000..cf4260d --- /dev/null +++ b/crates/npro-h1/src/fields.rs @@ -0,0 +1,83 @@ +//! What some header values mean, read as C reads them. + +use crate::table::HeaderTable; +use crate::token::Token; + +/// A `Content-Length` is not one. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct BadContentLength; + +impl core::fmt::Display for BadContentLength { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + f.write_str("bad Content-Length") + } +} + +impl core::error::Error for BadContentLength {} + +/// A `Content-Length` value: RFC 9110 8.6's `1*DIGIT`, with trailing +/// spaces tolerated as C has always tolerated them, and nothing that would +/// not fit in 64 bits. C's `lws_http_parse_content_length()`: no sign, no +/// leading space, no wrapping round. +/// +/// ``` +/// use npro_h1::fields::content_length; +/// +/// assert_eq!(content_length(b"10"), Ok(10)); +/// assert_eq!(content_length(b"10 "), Ok(10)); +/// assert!(content_length(b"10abc").is_err()); +/// assert!(content_length(b"-1").is_err()); +/// assert!(content_length(b"18446744073709551616").is_err()); +/// ``` +/// +/// # Errors +/// +/// [`BadContentLength`] for anything else. +pub fn content_length(value: &[u8]) -> Result<u64, BadContentLength> { + let digits = value.iter().take_while(|c| c.is_ascii_digit()).count(); + let (num, rest) = value.split_at(digits); + if num.is_empty() || rest.iter().any(|c| *c != b' ') { + return Err(BadContentLength); + } + num.iter().try_fold(0u64, |v, c| { + v.checked_mul(10) + .and_then(|v| v.checked_add(u64::from(c.wrapping_sub(b'0')))) + .ok_or(BadContentLength) + }) +} + +/// Whether a head's `Transfer-Encoding` is exactly one `chunked` coding, +/// the only one lws decodes: one instance of the header, whose value is +/// `chunked` in any case, with any spaces or tabs around it. A list of +/// codings would need each applied in turn, and a second instance of the +/// header makes a list however the sender split it. C's +/// `lws_http_te_is_chunked()`. +/// +/// ``` +/// use npro_h1::fields::transfer_encoding_is_chunked; +/// use npro_h1::head::{Config, Head, Side}; +/// +/// let mut h = Head::new([0u8; 512], Side::Client, Config::default())?; +/// h.rx(b"HTTP/1.1 200 OK\r\nTransfer-Encoding: Chunked \r\n\r\n")?; +/// assert!(transfer_encoding_is_chunked(h.table())); +/// # Ok::<(), Box<dyn core::error::Error>>(()) +/// ``` +#[must_use] +pub fn transfer_encoding_is_chunked<S: AsRef<[u8]> + AsMut<[u8]>>(table: &HeaderTable<S>) -> bool { + let mut f = table.fragments(Token::TransferEncoding); + let (Some(v), None) = (f.next(), f.next()) else { + return false; + }; + // C copies it into a 32 byte buffer, with room for its NUL + if v.is_empty() || v.len() > 30 { + return false; + } + let blank = |c: &u8| *c == b' ' || *c == b'\t'; + let start = v.iter().position(|c| !blank(c)).unwrap_or(v.len()); + let end = v + .iter() + .rposition(|c| !blank(c)) + .map_or(0, |n| n.saturating_add(1)); + v.get(start..end) + .is_some_and(|v| v.eq_ignore_ascii_case(b"chunked")) +} diff --git a/crates/npro-h1/src/head.rs b/crates/npro-h1/src/head.rs new file mode 100644 index 0000000..bc0d290 --- /dev/null +++ b/crates/npro-h1/src/head.rs @@ -0,0 +1,1224 @@ +//! An h1 head, parsed as it arrives: C's `lws_parse()`. +//! +//! A server's head is a request line and header lines; a client's is a +//! status line and header lines. Either ends at an empty line. The bytes +//! are taken one at a time, in pieces of any size, and nothing is kept +//! between pieces but [`Head`] itself, so a head split anywhere parses as +//! it does whole. +//! +//! What C's parser does is what this does, branch for branch, with C's +//! state in enums: `parser_state` is `State`, the URI decoder's `ues`, +//! `ups` and `post_literal_equal` are `Esc`, `Path` and `Arg`, and the +//! name matcher's `lextable_pos` and `unk_pos` are `Name`. Where C's +//! reasons for a refusal are labels (`bad_request_line`, `forbid`, +//! `too_large`...), here they are [`Cause`]s, and what a server answers +//! for one is [`Refused::answer`]. +//! +//! Porting it found three bugs in C, fixed there first (lws 3c8459075 and +//! the two before it), so this is C as it is now: a name's start is a +//! state of its own, where C had taken a record at offset 0 for no record +//! and lost nine bytes of every request's table; a strict server refuses a +//! header line starting with a bare CR; and a value's leading OWS is not +//! kept, a repeated header's included. + +use core::num::NonZeroU16; + +use crate::table::{CapacityTooLarge, Full, HeaderTable}; +use crate::token::{Lookup, Token, lookup}; + +const CR: u8 = b'\r'; +const LF: u8 = b'\n'; +const SP: u8 = b' '; +const HT: u8 = b'\t'; + +/// RFC 9112 2.2: a server ignores at least one empty line before a request +/// line. C ignores this many, and refuses the next. +const MAX_LEADING_EMPTY_LINES: u8 = 8; + +/// Whose head it is. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum Side { + /// A server, parsing a request. It is strict about line ends and + /// names, as what is in front of it may read them differently. + Server, + /// A client, parsing a response. It is tolerant of the servers it + /// talks to. + Client, +} + +/// What a server does with a request whose first token is no method it +/// knows. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +pub enum UnknownMethod { + /// Refuse it: 501 if the token ends at a SP, 400 if it is no method. + #[default] + Refuse, + /// Give the connection to another role, as C's + /// `LWS_SERVER_OPTION_FALLBACK_TO_APPLY_LISTEN_ACCEPT_CONFIG`: + /// [`Head::rx`] says [`Progress::Fallback`]. + Fallback, +} + +/// How a head is parsed, beyond the size of its table. +/// +/// ``` +/// use core::num::NonZeroU16; +/// use npro_h1::head::Config; +/// use npro_h1::token::Token; +/// +/// // as C's api-test-sansio context: a GET's target and a User-Agent of +/// // at most 33 and 16 bytes +/// let c = Config::new() +/// .with_limit(Token::GetUri, NonZeroU16::new(33).unwrap()) +/// .with_limit(Token::UserAgent, NonZeroU16::new(16).unwrap()); +/// assert_eq!(c.limit(Token::UserAgent), Some(16)); +/// assert_eq!(c.limit(Token::Host), None); +/// ``` +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct Config { + limits: [Option<NonZeroU16>; Token::COUNT], + unknown_method: UnknownMethod, +} + +impl Default for Config { + fn default() -> Self { + Self::new() + } +} + +impl Config { + /// No limit on any token but the table's size, and unknown methods + /// refused. + #[must_use] + pub const fn new() -> Self { + Self { + limits: [None; Token::COUNT], + unknown_method: UnknownMethod::Refuse, + } + } + + /// Limits `token`'s value, or each urlarg's for a method's target, to + /// `max` bytes: C's `token_limits`. A longer one refuses the head. + #[must_use] + pub fn with_limit(mut self, token: Token, max: NonZeroU16) -> Self { + if let Some(l) = self.limits.get_mut(token.index()) { + *l = Some(max); + } + self + } + + /// What a server does with an unknown method. + #[must_use] + pub const fn with_unknown_method(mut self, m: UnknownMethod) -> Self { + self.unknown_method = m; + self + } + + /// The most `token`'s value may have, if it is limited. + #[must_use] + pub fn limit(&self, token: Token) -> Option<u16> { + self.limits + .get(token.index()) + .copied() + .flatten() + .map(NonZeroU16::get) + } +} + +/// Why a head was refused. Each is one of C's ways out of `lws_parse()`. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum Cause { + /// A NUL byte, anywhere. + Nul, + /// A server's head has a LF without its CR. + BareLf, + /// A server's head has a CR without its LF. + BareCr, + /// A server's header name has this byte, which no field name may. + NameByte(u8), + /// A request line's method came again as a header name. + DuplicateMethod, + /// A server's header line is named for this token, which is no h1 + /// field name. + NotAFieldName(Token), + /// A server's head does not start with a request line. + NoRequestLine, + /// The request line ended after its target: HTTP/0.9, which lws does + /// not speak. + Http09, + /// More than eight empty lines came before the request line. + TooManyEmptyLines, + /// The request line's version is not `HTTP/` digit `.` digit. + BadVersion, + /// The request line's version is not 1.x. + VersionNotSupported, + /// The request line's method is one lws does not implement. + MethodNotImplemented, + /// The request target is refused. + Uri(UriFault), + /// The request target, or one of its urlargs, is longer than its + /// limit, or than the table holds. + UriTooLong, + /// Something else in the head is longer than its limit, or than the + /// table holds, or there were more pieces than it can track. + HeadTooLarge, +} + +/// What is wrong with a request target. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum UriFault { + /// A `%` followed by something not two hex digits. + BadEscape, + /// A `%` escape the target ended in the middle of. + UnfinishedEscape, + /// A control byte, or DEL, raw or `%` escaped. + ControlByte(u8), + /// The target starts with its `?`: there is no path. + EmptyPath, +} + +/// The status a server answers a refused head with. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum Answer { + /// 400. + BadRequest, + /// 403. + Forbidden, + /// 414. + UriTooLong, + /// 431. + HeaderFieldsTooLarge, + /// 501. + NotImplemented, + /// 505. + VersionNotSupported, +} + +impl Answer { + /// The status code. + #[must_use] + pub const fn code(self) -> u16 { + match self { + Self::BadRequest => 400, + Self::Forbidden => 403, + Self::UriTooLong => 414, + Self::HeaderFieldsTooLarge => 431, + Self::NotImplemented => 501, + Self::VersionNotSupported => 505, + } + } +} + +/// A refused head: why, and what a server says. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct Refused { + cause: Cause, + answer: Option<Answer>, +} + +impl Refused { + const fn new(side: Side, cause: Cause) -> Self { + let answer = match side { + Side::Client => None, + Side::Server => match cause { + Cause::Nul + | Cause::BareLf + | Cause::BareCr + | Cause::NameByte(_) + | Cause::DuplicateMethod + | Cause::NotAFieldName(_) => None, + Cause::NoRequestLine + | Cause::Http09 + | Cause::TooManyEmptyLines + | Cause::BadVersion => Some(Answer::BadRequest), + Cause::VersionNotSupported => Some(Answer::VersionNotSupported), + Cause::MethodNotImplemented => Some(Answer::NotImplemented), + Cause::Uri(_) => Some(Answer::Forbidden), + Cause::UriTooLong => Some(Answer::UriTooLong), + Cause::HeadTooLarge => Some(Answer::HeaderFieldsTooLarge), + }, + }; + Self { cause, answer } + } + + /// Why the head was refused. + #[must_use] + pub const fn cause(&self) -> Cause { + self.cause + } + + /// What a server answers before it closes: C's `LPR_REFUSED` and the + /// status it sends. `None` is C's `LPR_FAIL`, closing without a word, + /// and every refusal of a client's. + #[must_use] + pub const fn answer(&self) -> Option<Answer> { + self.answer + } +} + +impl core::fmt::Display for Refused { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + write!(f, "h1 head refused: {:?}", self.cause)?; + if let Some(a) = self.answer { + write!(f, ", answered {}", a.code())?; + } + Ok(()) + } +} + +impl core::error::Error for Refused {} + +/// How far a head got with the bytes it was given. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum Progress { + /// All of them were taken, and the head is not over. + More, + /// The head is over, at the end of the first `consumed` bytes. What + /// follows is a body, or the next head. + Complete { + /// How many of the bytes were the head's. + consumed: usize, + }, + /// The request is not one for http, and its connection goes to another + /// role, with every byte it sent: see [`UnknownMethod::Fallback`]. + Fallback, +} + +/// The version a server answers a request in: C's +/// `lws_h1_request_version()`. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum Version { + /// HTTP/1.0. + Http10, + /// HTTP/1.1, which a later 1.x is answered as (RFC 9110 2.5). + Http11, +} + +/// Where in a name the parser is: C's `lextable_pos` with `unk_pos`. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +enum Name { + /// Nothing of it yet. + Start, + /// Its bytes, from its record at `rec`, are the start of a spelling. + Matching { rec: u16 }, + /// It is no spelling, and is kept, from its record at `rec`. + Unknown { rec: u16 }, +} + +/// What the parser is in the middle of: C's `parser_state`. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +enum State { + /// A name: `WSI_TOKEN_NAME_PART`. + Name(Name), + /// A known token's value. + Value(Token), + /// An unknown header's value, its record at `rec` and the value at + /// `value`: `WSI_TOKEN_UNKNOWN_VALUE_PART`. + UnknownValue { rec: u16, value: u16 }, + /// The rest of a line not kept: `WSI_TOKEN_SKIPPING`. + Skipping, + /// A CR, which must be followed by LF: `WSI_TOKEN_SKIPPING_SAW_CR`. + SkippingSawCr, + /// The head is over: `WSI_PARSING_COMPLETE`. + Complete, + /// The head goes to the fallback role. + Fallback, + /// The head was refused. + Refused(Refused), +} + +/// A `%` escape in the target: C's `ues`. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +enum Esc { + Idle, + Percent, + /// The first hex digit, as its value. + PercentHigh(u8), +} + +/// Where the path is, for `//`, `/./` and `/../`: C's `ups`. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +enum Path { + Idle, + Slash, + SlashDot, + SlashDotDot, +} + +/// Which side of an urlarg's `=` the decoder is: C's +/// `post_literal_equal`. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +enum Arg { + Name, + Value, +} + +/// What a byte of the target becomes. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +enum Decoded { + /// This byte goes into the target. + Byte(u8), + /// Nothing goes in. + Swallow, +} + +/// What a byte did. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +enum Flow { + More, + Complete, + Fallback, +} + +/// An h1 head being parsed into its [`HeaderTable`]. +/// +/// ``` +/// use npro_h1::head::{Config, Head, Progress, Side}; +/// use npro_h1::token::Token; +/// +/// let mut h = Head::new([0u8; 1024], Side::Server, Config::default())?; +/// assert_eq!(h.rx(b"GET /a/../b?x=1 HTTP/1.1\r\nHo")?, Progress::More); +/// assert_eq!( +/// h.rx(b"st: example.com\r\n\r\nbody")?, +/// Progress::Complete { consumed: 19 } +/// ); +/// let t = h.table(); +/// assert_eq!(t.first(Token::GetUri), Some(&b"/b"[..])); +/// assert_eq!(t.first(Token::UriArgs), Some(&b"x=1"[..])); +/// assert_eq!(t.first(Token::Host), Some(&b"example.com"[..])); +/// # Ok::<(), Box<dyn core::error::Error>>(()) +/// ``` +#[derive(Clone, Debug)] +pub struct Head<S> { + table: HeaderTable<S>, + side: Side, + config: Config, + state: State, + esc: Esc, + path: Path, + arg: Arg, + /// The limit of the value being filled: C's `current_token_limit`. + limit: Option<u16>, + empty_lines: u8, +} + +impl<S: AsRef<[u8]> + AsMut<[u8]>> Head<S> { + /// A head to parse into a table in `storage`. + /// + /// # Errors + /// + /// [`CapacityTooLarge`] if `storage` is past the most a table holds. + pub fn new(storage: S, side: Side, config: Config) -> Result<Self, CapacityTooLarge> { + Ok(Self::with_table(HeaderTable::new(storage)?, side, config)) + } + + /// A head to parse into `table`, after what it holds: how a client + /// parses its response into the table holding its own request. + #[must_use] + pub const fn with_table(table: HeaderTable<S>, side: Side, config: Config) -> Self { + Self { + table, + side, + config, + state: State::Name(Name::Start), + esc: Esc::Idle, + path: Path::Idle, + arg: Arg::Name, + limit: None, + empty_lines: 0, + } + } + + /// Readies the head for the next one, with the table emptied: C's + /// `lws_header_table_reset()`. + pub fn reset(&mut self) { + self.table.reset(); + self.restart(); + } + + /// Readies the head for another, dropping from the table what was put + /// there since its [`snapshot`](HeaderTable::snapshot): how a client + /// goes on to the final response after an interim one, as C's + /// `lws_header_table_rx_rewind()`. + pub fn rewind(&mut self) { + self.table.rewind(); + self.restart(); + } + + const fn restart(&mut self) { + self.state = State::Name(Name::Start); + self.esc = Esc::Idle; + self.path = Path::Idle; + self.arg = Arg::Name; + self.limit = None; + self.empty_lines = 0; + } + + /// The table, with what the head has put in it so far. + #[must_use] + pub const fn table(&self) -> &HeaderTable<S> { + &self.table + } + + /// The table, for its owner to add to, as a client does with its own + /// request before it parses the response. + #[must_use] + pub const fn table_mut(&mut self) -> &mut HeaderTable<S> { + &mut self.table + } + + /// Whose head it is. + #[must_use] + pub const fn side(&self) -> Side { + self.side + } + + /// Whether the head is over. + #[must_use] + pub const fn is_complete(&self) -> bool { + matches!(self.state, State::Complete) + } + + /// The version a server answers in, from the request line: HTTP/1.1 + /// for `HTTP/1.1` to `HTTP/1.9`, otherwise HTTP/1.0. C's + /// `lws_h1_request_version()`. + #[must_use] + pub fn request_version(&self) -> Version { + // as C copies it, into a buffer of 11 with room for its NUL + let mut v = [0u8; 10]; + let total = self.table.copy(Token::Http, &mut v).unwrap_or(0); + if total > 7 + && v.get(5) == Some(&b'1') + && v.get(7).is_some_and(|d| (b'1'..=b'9').contains(d)) + { + Version::Http11 + } else { + Version::Http10 + } + } + + /// Takes the next bytes of the head. + /// + /// # Errors + /// + /// [`Refused`], as soon as a byte makes the head one lws refuses. The + /// head stays refused: every later call says the same. + pub fn rx(&mut self, bytes: &[u8]) -> Result<Progress, Refused> { + match self.state { + State::Refused(r) => return Err(r), + State::Complete => return Ok(Progress::Complete { consumed: 0 }), + State::Fallback => return Ok(Progress::Fallback), + State::Name(_) + | State::Value(_) + | State::UnknownValue { .. } + | State::Skipping + | State::SkippingSawCr => {} + } + for (i, &c) in bytes.iter().enumerate() { + match self.byte(c) { + Ok(Flow::More) => {} + Ok(Flow::Complete) => { + self.state = State::Complete; + return Ok(Progress::Complete { + consumed: i.saturating_add(1), + }); + } + Ok(Flow::Fallback) => { + self.state = State::Fallback; + return Ok(Progress::Fallback); + } + Err(cause) => { + let r = Refused::new(self.side, cause); + self.state = State::Refused(r); + return Err(r); + } + } + } + Ok(Progress::More) + } + + /// Whether a server's request line has had its method. + fn method_seen(&self) -> bool { + Token::METHODS.iter().any(|m| self.table.is_present(*m)) + } + + /// A server, strict about line ends and names: C's + /// `lws_h1_srv_strict()`. + fn strict(&self) -> bool { + self.side == Side::Server + } + + /// A server still waiting for its request line: C's + /// `lws_h1_srv_awaits_request_line()`. + fn awaits_request_line(&self) -> bool { + self.strict() && !self.method_seen() + } + + /// A byte a server refuses in a header name, once it has had its + /// request line: C's `lws_h1_srv_bad_name_char()`. + fn bad_name_byte(&self, c: u8) -> bool { + self.strict() && c != b':' && !name_byte(c) && self.method_seen() + } + + /// Whether `t`, matched as a name, can be taken here: C's + /// `lws_h1_token_usable()`. + fn usable(&self, t: Token) -> bool { + if t == Token::Challenge || t.spelling().last() == Some(&b':') { + return true; + } + if self.side == Side::Client { + return t == Token::Http || t == Token::Http10; + } + !self.method_seen() && t.is_method() + } + + /// What is too large where the parser is: the target, or the rest. + const fn too_large(t: Token) -> Cause { + if matches!( + t, + Token::GetUri + | Token::PostUri + | Token::OptionsUri + | Token::PutUri + | Token::PatchUri + | Token::DeleteUri + | Token::Connect + | Token::HeadUri + ) { + Cause::UriTooLong + } else { + Cause::HeadTooLarge + } + } + + /// One byte of the head. + fn byte(&mut self, c: u8) -> Result<Flow, Cause> { + if c == 0 { + return Err(Cause::Nul); + } + match self.state { + State::UnknownValue { rec, value } => self.unknown_value(c, rec, value), + State::Value(t) => self.value(c, t), + State::Name(n) => self.name(c, n), + State::Skipping => { + if c == LF { + if self.strict() { + return Err(Cause::BareLf); + } + self.state = State::Name(Name::Start); + } + if c == CR { + self.state = State::SkippingSawCr; + } + Ok(Flow::More) + } + State::SkippingSawCr => { + if self.esc != Esc::Idle { + return Err(Cause::Uri(UriFault::UnfinishedEscape)); + } + if c == LF { + self.state = State::Name(Name::Start); + return Ok(Flow::More); + } + if self.strict() { + return Err(Cause::BareCr); + } + self.state = State::Skipping; + Ok(Flow::More) + } + State::Complete | State::Fallback | State::Refused(_) => Ok(Flow::More), + } + } + + fn unknown_value(&mut self, c: u8, rec: u16, value: u16) -> Result<Flow, Cause> { + // the value ends at the CR, whose LF is then checked for like any + // other header's + if c == LF && self.strict() { + return Err(Cause::BareLf); + } + if c == CR || c == LF { + self.table.end_record_value(rec, value); + self.state = if c == CR { + State::SkippingSawCr + } else { + State::Name(Name::Start) + }; + return Ok(Flow::More); + } + // without its leading whitespace + if self.table.pos() != value || (c != SP && c != HT) { + self.table.push(c).map_err(|_| Cause::HeadTooLarge)?; + } + Ok(Flow::More) + } + + fn value(&mut self, mut c: u8, t: Token) -> Result<Flow, Cause> { + if t.is_method() { + // extra SP between the method and the target + if self.table.first_len(t) == 0 && c == SP { + return Ok(Flow::More); + } + if c == CR || c == LF { + return Err(Cause::Http09); + } + if c == SP { + return self.end_of_target(); + } + match self.urldecode(c)? { + Decoded::Swallow => return Ok(Flow::More), + Decoded::Byte(d) => c = d, + } + } else { + // the OWS before a header's value is not part of it (RFC 9110 + // 5.5): swallowed while the piece has nothing of the value's + // own, a repeated header's only the SP joining it to the one + // before. The first line's version is no header: a HT is no + // SP there + let ows = c == SP || (c == HT && t != Token::Http && t != Token::Http10); + let joining = u16::from(!self.table.filling_first(t)); + if ows && self.table.current_len() == joining { + return Ok(Flow::More); + } + } + // the end of the line + if t != Token::Challenge && (c == CR || c == LF) { + if self.esc != Esc::Idle { + return Err(Cause::Uri(UriFault::UnfinishedEscape)); + } + if t == Token::Http && self.side == Side::Server { + if let Some(refusal) = version_refusal(self.table.current()) { + return Err(refusal); + } + } + if c == LF { + if self.strict() { + return Err(Cause::BareLf); + } + self.state = State::Name(Name::Start); + } else { + self.state = State::SkippingSawCr; + } + c = 0; + } + self.issue(c, t)?; + if c == 0 { + // the NUL ending the value is not part of it + self.table.uncount(); + } + Ok(Flow::More) + } + + /// Adds `c` to the value of `t` being filled. + fn issue(&mut self, c: u8, t: Token) -> Result<(), Cause> { + self.table + .issue(c, self.limit) + .map_err(|_| Self::too_large(t)) + } + + /// The SP ending a request target. + fn end_of_target(&mut self) -> Result<Flow, Cause> { + // a target starts with a /, but only while it is the path: after + // a ?, an empty query stays empty + if !self.table.is_present(Token::UriArgs) && self.table.current_len() == 0 { + self.issue(b'/', Token::GetUri)?; + } + if self.path == Path::SlashDotDot { + self.table.back_up_a_segment(); + } + self.issue(0, Token::GetUri)?; + self.table.uncount(); + // the version, under the target's limit, as C has it + self.state = State::Value(Token::Http); + self.start_fragment(Token::Http) + } + + /// A value of `t` begins. + fn start_fragment(&mut self, t: Token) -> Result<Flow, Cause> { + let chained = self + .table + .start_fragment(t) + .map_err(|_| Self::too_large(t))?; + if chained { + self.issue(SP, t)?; + } + Ok(Flow::More) + } + + fn name(&mut self, c: u8, n: Name) -> Result<Flow, Cause> { + if n == Name::Start && c == LF { + if self.strict() { + return Err(Cause::BareLf); + } + // a broken peer's empty line + return self.complete(); + } + // an empty line where the request line should start, before + // anything of the head: skipped, a few of them + if c == CR && n == Name::Start && !self.table.has_fragments() && self.awaits_request_line() + { + self.empty_lines = self.empty_lines.saturating_add(1); + if self.empty_lines > MAX_LEADING_EMPTY_LINES { + return Err(Cause::TooManyEmptyLines); + } + self.state = State::SkippingSawCr; + return Ok(Flow::More); + } + // a field name is not empty: no ':' starts one + if c == b':' && n == Name::Start && self.strict() && self.method_seen() { + return Err(Cause::NameByte(c)); + } + let c = c.to_ascii_lowercase(); + // in case it is a header lws does not know, the name is kept as it + // comes, and dropped if it turns out to be one it knows + let rec = match n { + Name::Start => self.table.begin_record(), + Name::Matching { rec } | Name::Unknown { rec } => rec, + }; + self.table.push(c).map_err(|_| Cause::HeadTooLarge)?; + + if let Name::Unknown { .. } = n { + return self.unknown_name(c, rec); + } + match lookup(self.table.record_name(rec)) { + Lookup::Prefix => { + self.state = State::Name(Name::Matching { rec }); + Ok(Flow::More) + } + Lookup::Nothing => self.no_match(c, rec), + Lookup::Matched(t) => self.matched(t, rec), + } + } + + /// The next byte of a name lws does not know. + fn unknown_name(&mut self, c: u8, rec: u16) -> Result<Flow, Cause> { + if self.awaits_request_line() { + // a first token lws does not know: a method, if it ends at SP + if c == SP { + return Err(Cause::MethodNotImplemented); + } + if c == b':' || !name_byte(c) { + return Err(Cause::NoRequestLine); + } + self.state = State::Name(Name::Unknown { rec }); + return Ok(Flow::More); + } + if self.bad_name_byte(c) { + return Err(Cause::NameByte(c)); + } + if c == b':' { + return Ok(self.unknown_name_ended(rec)); + } + self.state = State::Name(Name::Unknown { rec }); + Ok(Flow::More) + } + + /// The name, ending in `c`, is the start of no spelling. + fn no_match(&mut self, c: u8, rec: u16) -> Result<Flow, Cause> { + if self.side == Side::Client { + // a client keeps a header it does not know, to its ':' + if c == b':' { + return Ok(self.unknown_name_ended(rec)); + } + self.state = State::Name(Name::Unknown { rec }); + return Ok(Flow::More); + } + if self.method_seen() { + // the bytes before c matched the start of a token, which need + // not be a field name's: the CR of "\r\n", which then started + // the line, is a bare CR + if let Some((_, before)) = self.table.record_name(rec).split_last() { + if let Some(b) = before.iter().find(|b| !name_byte(**b)) { + return Err(if *b == CR { + Cause::BareCr + } else { + Cause::NameByte(*b) + }); + } + } + // c is where the name stopped matching any lws knows: eg the + // SP of "host :", or the ':' of "accept-lang:" + if self.bad_name_byte(c) { + return Err(Cause::NameByte(c)); + } + if c == b':' { + return Ok(self.unknown_name_ended(rec)); + } + self.state = State::Name(Name::Unknown { rec }); + return Ok(Flow::More); + } + // a method lws does not know, or no request line at all + if self.config.unknown_method == UnknownMethod::Fallback { + return Ok(Flow::Fallback); + } + if c == SP { + return Err(Cause::MethodNotImplemented); + } + if c == b':' || !name_byte(c) { + return Err(Cause::NoRequestLine); + } + self.state = State::Name(Name::Unknown { rec }); + Ok(Flow::More) + } + + /// The name, from its record at `rec`, is the whole spelling of `t`. + fn matched(&mut self, t: Token, rec: u16) -> Result<Flow, Cause> { + if t.is_method() && self.table.is_present(t) { + return Err(Cause::DuplicateMethod); + } + if !self.usable(t) { + // a server's head starts with its request line... + if self.awaits_request_line() { + return Err(Cause::NoRequestLine); + } + // ...and a server refuses a name with a ':' or SP in it, as + // any other; a client skips the line + if !t.spelling().iter().all(|b| name_byte(*b)) { + if self.strict() { + return Err(Cause::NotAFieldName(t)); + } + self.table.truncate(rec); + self.state = State::Skipping; + return Ok(Flow::More); + } + // otherwise it is a name lws does not know, kept from its + // first byte, and going on to its ':' + self.state = State::Name(Name::Unknown { rec }); + return Ok(Flow::More); + } + // a header lws knows: the name kept for it is dropped + self.table.truncate(rec); + // Sec-WebSocket-Origin is Origin, as JWebSocket sends it + let t = if t == Token::WsOrigin { + Token::Origin + } else { + t + }; + self.state = State::Value(t); + self.path = Path::Idle; + self.limit = self.config.limit(t); + if t == Token::Challenge { + return self.complete(); + } + self.start_fragment(t) + } + + /// The ':' ending the name of a header lws does not know. + fn unknown_name_ended(&mut self, rec: u16) -> Flow { + self.table.end_record_name(rec); + self.state = State::UnknownValue { + rec, + value: self.table.pos(), + }; + Flow::More + } + + /// The empty line ending the head. + fn complete(&mut self) -> Result<Flow, Cause> { + if self.esc != Esc::Idle { + return Err(Cause::Uri(UriFault::UnfinishedEscape)); + } + // a server's head starts with its request line + if self.awaits_request_line() { + return Err(Cause::NoRequestLine); + } + Ok(Flow::Complete) + } + + /// A byte of a request target: `%` escapes, then the control bytes + /// refused, then `//`, `/./` and `/../` taken out, never above the + /// root, then the urlargs split: C's `lws_parse_urldecode()`. + fn urldecode(&mut self, c: u8) -> Result<Decoded, Cause> { + let mut c = c; + let mut escaped = false; + match self.esc { + Esc::Idle => { + if c == b'%' { + self.esc = Esc::Percent; + return Ok(Decoded::Swallow); + } + } + Esc::Percent => { + let h = hex(c).ok_or(Cause::Uri(UriFault::BadEscape))?; + self.esc = Esc::PercentHigh(h); + return Ok(Decoded::Swallow); + } + Esc::PercentHigh(h) => { + let l = hex(c).ok_or(Cause::Uri(UriFault::BadEscape))?; + c = (h << 4) | l; + escaped = true; + self.esc = Esc::Idle; + } + } + + // no control byte or DEL in a target, raw or escaped + if c < 0x20 || c == 0x7f { + return Err(Cause::Uri(UriFault::ControlByte(c))); + } + self.unescaped(c, escaped) + } + + /// A byte of a request target, `%` decoded if `escaped`. + fn unescaped(&mut self, mut c: u8, escaped: bool) -> Result<Decoded, Cause> { + let t = Token::GetUri; + let args = self.table.is_present(Token::UriArgs); + match self.path { + Path::Idle => { + // the urlargs' separators, once a ? has started them + if (c == b'&' || c == b';') && !escaped && args { + self.issue(0, t)?; + self.table.uncount(); + self.table.next_arg().map_err(|_| Cause::UriTooLong)?; + self.arg = Arg::Name; + return Ok(Decoded::Swallow); + } + // an escaped = is not the one ending an urlarg's name + if c == b'=' && escaped && args && self.arg == Arg::Name { + c = b'_'; + } + if c == b'=' && !escaped { + self.arg = Arg::Value; + } + // + is a space in the query, form encoding's; in the path + // it is a + + if c == b'+' && !escaped && args { + c = SP; + } + if c == b'/' && !args { + self.path = Path::Slash; + } + } + Path::Slash => { + if c == b'/' { + return Ok(Decoded::Swallow); + } + if c == b'.' { + self.path = Path::SlashDot; + return Ok(Decoded::Swallow); + } + self.path = Path::Idle; + } + Path::SlashDot => { + if c == b'.' { + self.path = Path::SlashDotDot; + return Ok(Decoded::Swallow); + } + if c == b'/' { + self.path = Path::Slash; + return Ok(Decoded::Swallow); + } + if c == b'?' && !escaped { + // "/.?" ends the path in "/.", as "/." at its end + self.path = Path::Slash; + } else { + // "/.dir": the . was part of it + self.path = Path::Idle; + self.issue(b'.', t)?; + } + } + Path::SlashDotDot => { + if c == b'/' || c == b'?' { + self.table.back_up_a_segment(); + self.path = Path::Slash; + // the / backed up to stands for a / here, but a ? goes + // on to start the urlargs + if self.table.current_len() <= 1 && c != b'?' { + return Ok(Decoded::Swallow); + } + } else { + // "/..x": the dots were part of it + self.issue(b'.', t)?; + self.issue(b'.', t)?; + self.path = Path::Idle; + } + } + } + + if c == b'?' && !escaped && !args { + if self.esc != Esc::Idle { + return Err(Cause::Uri(UriFault::UnfinishedEscape)); + } + // no target is all query + if self.table.current_len() == 0 { + return Err(Cause::Uri(UriFault::EmptyPath)); + } + self.issue(0, t)?; + self.table.uncount(); + self.table + .start_args() + .map_err(|_: Full| Cause::UriTooLong)?; + self.arg = Arg::Name; + self.path = Path::Idle; + return Ok(Decoded::Swallow); + } + + Ok(Decoded::Byte(c)) + } +} + +/// Whether a server takes `c` in a header name: C's +/// `lws_http_field_name_char_valid()`, past the first byte. No controls, +/// SP, DEL, non-ASCII, uppercase or ':'. +const fn name_byte(c: u8) -> bool { + c > 0x20 && c < 0x7f && !c.is_ascii_uppercase() && c != b':' +} + +const fn hex(c: u8) -> Option<u8> { + match c { + b'0'..=b'9' => Some(c.wrapping_sub(b'0')), + b'a'..=b'f' => Some(c.wrapping_sub(b'a').wrapping_add(10)), + b'A'..=b'F' => Some(c.wrapping_sub(b'A').wrapping_add(10)), + _ => None, + } +} + +/// RFC 9112 2.3: `HTTP/` digit `.` digit, which a server speaks if the +/// major version is 1: C's `lws_h1_version_refusal()`. +const fn version_refusal(v: &[u8]) -> Option<Cause> { + let [b'H', b'T', b'T', b'P', b'/', major, b'.', minor] = *v else { + return Some(Cause::BadVersion); + }; + if !major.is_ascii_digit() || !minor.is_ascii_digit() { + return Some(Cause::BadVersion); + } + if major != b'1' { + return Some(Cause::VersionNotSupported); + } + None +} + +#[cfg(test)] +mod tests { + use super::*; + + fn server(head: &[u8]) -> (Result<Progress, Refused>, Head<[u8; 512]>) { + let mut h = Head::new([0u8; 512], Side::Server, Config::new()).unwrap(); + let r = h.rx(head); + (r, h) + } + + #[test] + fn a_server_head_starts_the_table_at_its_first_byte() { + // the first name's record is at 0, and begun once + let (r, h) = server(b"GET / HTTP/1.1\r\n\r\n"); + assert_eq!(r, Ok(Progress::Complete { consumed: 18 })); + // "/" and its NUL, "HTTP/1.1" and its NUL + assert_eq!(h.table().used(), 2 + 9); + } + + #[test] + fn a_lf_second_is_a_name_byte_not_a_line_end() { + // a first token no method has, not the end of an empty head + let (r, _) = server(b"G\nET / HTTP/1.1\r\n\r\n"); + assert_eq!(r.map_err(|e| e.cause()), Err(Cause::NoRequestLine)); + assert_eq!(r.err().and_then(|e| e.answer()), Some(Answer::BadRequest)); + } + + #[test] + fn a_cr_starting_a_header_line_is_a_bare_cr() { + let (r, _) = server(b"GET / HTTP/1.1\r\n\rX-A: b\r\n\r\n"); + assert_eq!(r.map_err(|e| e.cause()), Err(Cause::BareCr)); + assert_eq!(r.err().and_then(|e| e.answer()), None); + // before the name of a header lws knows too + let (known, _) = server(b"GET / HTTP/1.1\r\n\rHost: b\r\n\r\n"); + assert_eq!(known.map_err(|e| e.cause()), Err(Cause::BareCr)); + // a client is tolerant, as C is: the CR is the name's + let mut h = Head::new([0u8; 512], Side::Client, Config::new()).unwrap(); + assert!(h.rx(b"HTTP/1.1 200 OK\r\n\rX-A: b\r\n\r\n").is_ok()); + assert_eq!(h.table().unknown(b"\rx-a:"), Some(&b"b"[..])); + } + + #[test] + fn a_header_named_like_a_token_it_cannot_be_is_kept_whole() { + let (r, h) = server(b"GET / HTTP/1.1\r\nuri-args: x\r\nPutX: y\r\n\r\n"); + assert!(r.is_ok()); + assert_eq!(h.table().unknown(b"uri-args:"), Some(&b"x"[..])); + assert_eq!(h.table().unknown(b"putx:"), Some(&b"y"[..])); + assert!(!h.table().is_present(Token::UriArgs)); + } + + #[test] + fn a_values_leading_ows_is_not_kept() { + // RFC 9110 5.5, repeated or not: a repeated header's piece keeps + // only the SP joining it to the one before + let (_, h) = server(b"GET / HTTP/1.1\r\nAccept: \t a\r\nAccept:\t b \r\n\r\n"); + let v: [&[u8]; 2] = [b"a", b" b "]; + assert!(h.table().fragments(Token::Accept).eq(v)); + let mut joined = [0u8; 8]; + assert_eq!(h.table().copy(Token::Accept, &mut joined), Ok(5)); + assert_eq!(&joined[..5], b"a, b "); + // the version is no header: a HT before it is no SP + let (r, _) = server(b"GET /\tHTTP/1.1\r\n\r\n"); + assert!(r.is_err()); + } + + #[test] + fn the_version_is_held_to_the_targets_limit() { + // as C: the limit taken at the method stays for the version + let cfg = Config::new().with_limit(Token::GetUri, NonZeroU16::new(4).unwrap()); + let mut h = Head::new([0u8; 512], Side::Server, cfg).unwrap(); + let r = h.rx(b"GET /abc HTTP/1.1\r\n\r\n"); + assert_eq!(r.map_err(|e| e.cause()), Err(Cause::HeadTooLarge)); + assert_eq!( + r.err().and_then(|e| e.answer()), + Some(Answer::HeaderFieldsTooLarge) + ); + let mut longer = Head::new([0u8; 512], Side::Server, cfg).unwrap(); + let refused = longer.rx(b"GET /abcd HTTP/1.1\r\n\r\n"); + assert_eq!( + refused.err().and_then(|e| e.answer()), + Some(Answer::UriTooLong) + ); + } + + #[test] + fn a_refused_head_stays_refused_and_a_complete_one_takes_no_more() { + let (r, mut h) = server(b"GET / HTTP/2.0\r\n"); + let e = r.unwrap_err(); + assert_eq!(e.answer(), Some(Answer::VersionNotSupported)); + assert_eq!(h.rx(b"\r\n"), Err(e)); + let (_, mut done) = server(b"GET / HTTP/1.1\r\n\r\n"); + assert_eq!(done.rx(b"GET"), Ok(Progress::Complete { consumed: 0 })); + } + + #[test] + fn the_version_answered_in() { + for (v, want) in [ + (&b"HTTP/1.0"[..], Version::Http10), + (b"HTTP/1.1", Version::Http11), + (b"HTTP/1.9", Version::Http11), + ] { + let mut head = b"GET / ".to_vec(); + head.extend_from_slice(v); + head.extend_from_slice(b"\r\n\r\n"); + let (r, h) = server(&head); + assert!(r.is_ok()); + assert_eq!(h.request_version(), want); + } + } + + #[test] + fn a_fallback_role_takes_an_unknown_method() { + let cfg = Config::new().with_unknown_method(UnknownMethod::Fallback); + let mut h = Head::new([0u8; 512], Side::Server, cfg).unwrap(); + assert_eq!(h.rx(b"SSH-2.0\r\n"), Ok(Progress::Fallback)); + assert_eq!(h.rx(b"x"), Ok(Progress::Fallback)); + } + + #[test] + fn an_interim_response_is_rewound_to_the_clients_own_request() { + let mut t = HeaderTable::new([0u8; 512]).unwrap(); + t.create(Token::ClientUri, b"/x").unwrap(); + t.snapshot(); + let mut h = Head::with_table(t, Side::Client, Config::new()); + let r = h.rx(b"HTTP/1.1 100 Continue\r\n\r\nHTTP/1.1 200 OK\r\n\r\n"); + assert_eq!(r, Ok(Progress::Complete { consumed: 25 })); + h.rewind(); + assert!(!h.table().is_present(Token::Http)); + assert_eq!( + h.rx(b"HTTP/1.1 200 OK\r\n\r\n"), + Ok(Progress::Complete { consumed: 19 }) + ); + assert_eq!(h.table().first(Token::Http), Some(&b"200 OK"[..])); + assert_eq!(h.table().first(Token::ClientUri), Some(&b"/x"[..])); + } +} diff --git a/crates/npro-h1/src/lib.rs b/crates/npro-h1/src/lib.rs new file mode 100644 index 0000000..9dd7392 --- /dev/null +++ b/crates/npro-h1/src/lib.rs @@ -0,0 +1,21 @@ +//! npro-h1: h1 for npro, sans-IO. +//! +//! The port of C libwebsockets' h1 parsing (`lib/sansio/http/parsers.c`): +//! +//! - [`head`]: a request or response head, parsed as it arrives, with C's +//! limits and refusals, into +//! - [`table`]: where the head's headers are kept, in caller-owned storage, +//! found by +//! - [`token`]: the headers lws knows by name; +//! - [`fields`]: what a few header values mean, as C reads them. +//! +//! The h1 client and server transactions come next, in phase 1d of the +//! port plan. + +#![no_std] +#![forbid(unsafe_code)] + +pub mod fields; +pub mod head; +pub mod table; +pub mod token; diff --git a/crates/npro-h1/src/table.rs b/crates/npro-h1/src/table.rs new file mode 100644 index 0000000..408d86a --- /dev/null +++ b/crates/npro-h1/src/table.rs @@ -0,0 +1,788 @@ +//! Where a head's headers are kept: C's `struct allocated_headers`, the +//! "ah". +//! +//! The values are laid down in one caller-owned buffer as they arrive, each +//! piece a fragment: where it starts and how long it is. A token's first +//! fragment is found from the token, and a header that comes again chains +//! another fragment onto it. A header lws does not know is kept as a +//! record in the same buffer, its name and value behind eight bytes saying +//! how long each is and where the next record is, the records linked in +//! the order their names ended. +//! +//! The layout is C's, byte for byte in what it uses up: each value ends in +//! a NUL C's string users need, a record starts with C's eight bytes, and +//! the fragment slots are C's `WSI_TOKEN_COUNT`. So a head fills the table +//! at the same byte as it fills C's, and is refused there as C refuses it. +//! +//! Where C says a token is present by a nonzero fragment index, and a link +//! ends at a zero one, here they are `Slot` and an `Option`. + +use crate::token::Token; + +/// The most a table holds, as C caps `max_http_header_data`. +pub const MAX_CAPACITY: usize = 32768; + +/// What C's `max_http_header_data` is unless it is set. +pub const DEFAULT_CAPACITY: usize = 4096; + +/// The bytes of an unknown header's record before its name: its name's +/// length, its value's, and the next record's offset, big endian, as C's +/// `UHO_NLEN`, `UHO_VLEN` and `UHO_LL`. +const RECORD: u16 = 8; + +/// A table's storage is past [`MAX_CAPACITY`]. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct CapacityTooLarge; + +impl core::fmt::Display for CapacityTooLarge { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + write!(f, "a header table holds at most {MAX_CAPACITY} bytes") + } +} + +impl core::error::Error for CapacityTooLarge {} + +/// What was asked to go into the table does not fit: its storage is full, +/// or every fragment slot is taken. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct Full; + +impl core::fmt::Display for Full { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + f.write_str("the header table is full") + } +} + +impl core::error::Error for Full {} + +/// A buffer is too small for what was to be copied into it. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct TooSmall { + /// How many bytes it needs. + pub needed: usize, +} + +impl core::fmt::Display for TooSmall { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + write!( + f, + "the buffer is too small: {} bytes are needed", + self.needed + ) + } +} + +impl core::error::Error for TooSmall {} + +/// A fragment slot, `1..Token::COUNT` as C numbers `ah->frags[]`. +#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord)] +pub(crate) struct FragId(u8); + +impl FragId { + fn slot(self) -> usize { + usize::from(self.0) + } +} + +/// One piece of a value: C's `struct lws_fragments`. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +struct Frag { + offset: u16, + len: u16, + /// The fragment continuing this token's value, if any. + next: Option<FragId>, +} + +/// Whether a token is in the table, and where its value starts. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +enum Slot { + #[default] + Absent, + Present(FragId), +} + +/// The unknown headers' records, as C's `unk_ll_head` and `unk_ll_tail`. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +enum Unknowns { + #[default] + None, + Linked { + head: u16, + tail: u16, + }, +} + +/// Where a client's response starts in the table, as C's `rx_snap_*`. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +struct Mark { + pos: u16, + nfrag: u8, + unknowns: Unknowns, +} + +/// A head's headers. +/// +/// `S` is the storage: `[u8; 4096]` on a device, a `Box<[u8]>` where there +/// is an allocator. Its length is the table's capacity, C's +/// `max_http_header_data`, at most [`MAX_CAPACITY`]. +/// +/// [`crate::head::Head`] fills one from a peer's bytes. Values are bytes, +/// not text: a header's value may be anything but NUL, CR and LF. +/// +/// ``` +/// use npro_h1::table::HeaderTable; +/// use npro_h1::token::Token; +/// +/// let mut t = HeaderTable::new([0u8; 256])?; +/// t.create(Token::ClientUri, b"/x")?; +/// assert_eq!(t.first(Token::ClientUri), Some(&b"/x"[..])); +/// assert!(!t.is_present(Token::Host)); +/// # Ok::<(), Box<dyn core::error::Error>>(()) +/// ``` +#[derive(Clone, Debug)] +pub struct HeaderTable<S> { + data: S, + capacity: u16, + pos: u16, + nfrag: u8, + frags: [Frag; Token::COUNT], + index: [Slot; Token::COUNT], + unknowns: Unknowns, + mark: Mark, +} + +impl<S: AsRef<[u8]> + AsMut<[u8]>> HeaderTable<S> { + /// An empty table in `data`. + /// + /// # Errors + /// + /// [`CapacityTooLarge`] if `data` is longer than [`MAX_CAPACITY`]. + pub fn new(data: S) -> Result<Self, CapacityTooLarge> { + if data.as_ref().len() > MAX_CAPACITY { + return Err(CapacityTooLarge); + } + let capacity = u16::try_from(data.as_ref().len()).map_err(|_| CapacityTooLarge)?; + Ok(Self { + data, + capacity, + pos: 0, + nfrag: 0, + frags: [Frag::default(); Token::COUNT], + index: [Slot::Absent; Token::COUNT], + unknowns: Unknowns::None, + mark: Mark::default(), + }) + } + + /// Empties the table, as C's `_lws_header_table_reset()`. + pub fn reset(&mut self) { + self.pos = 0; + self.nfrag = 0; + self.frags = [Frag::default(); Token::COUNT]; + self.index = [Slot::Absent; Token::COUNT]; + self.unknowns = Unknowns::None; + self.mark = Mark::default(); + } + + /// How many bytes the table holds. + #[must_use] + pub fn capacity(&self) -> usize { + usize::from(self.capacity) + } + + /// How many bytes are used. + #[must_use] + pub fn used(&self) -> usize { + usize::from(self.pos) + } + + /// Whether `token` is in the table: C's `lws_hdr_extant()`. + #[must_use] + pub fn is_present(&self, token: Token) -> bool { + matches!(self.slot(token), Slot::Present(_)) + } + + /// The pieces of `token`'s value, one for each time the header came, or + /// for the urlargs one per argument. Empty if the token is absent. + /// + /// ``` + /// use npro_h1::head::{Config, Head, Side}; + /// use npro_h1::token::Token; + /// + /// let mut h = Head::new([0u8; 512], Side::Server, Config::default())?; + /// h.rx(b"GET /?a=1&b=2 HTTP/1.1\r\n\r\n")?; + /// let args: Vec<&[u8]> = h.table().fragments(Token::UriArgs).collect(); + /// assert_eq!(args, [&b"a=1"[..], b"b=2"]); + /// # Ok::<(), Box<dyn core::error::Error>>(()) + /// ``` + #[must_use] + pub fn fragments(&self, token: Token) -> Fragments<'_> { + Fragments { + data: self.data.as_ref(), + frags: &self.frags, + next: match self.slot(token) { + Slot::Present(f) => Some(f), + Slot::Absent => None, + }, + } + } + + /// The first piece of `token`'s value: C's `lws_hdr_simple_ptr()`. + #[must_use] + pub fn first(&self, token: Token) -> Option<&[u8]> { + self.fragments(token).next() + } + + /// How long `token`'s value is with its pieces joined, as + /// [`copy`](Self::copy) joins them: C's `lws_hdr_total_length()`. + #[must_use] + pub fn total_len(&self, token: Token) -> usize { + let mut n = 0usize; + for (i, f) in self.fragments(token).enumerate() { + if i > 0 { + n = n.saturating_add(1); + } + n = n.saturating_add(f.len()); + } + n + } + + /// Copies `token`'s value into `out`, its pieces joined by `,`, or by + /// `;` for cookies and `&` for urlargs: C's `lws_hdr_copy()`. Returns + /// how much of `out` it used, which is nothing if the token is absent. + /// + /// # Errors + /// + /// [`TooSmall`], copying nothing, if the value does not fit. + pub fn copy(&self, token: Token, out: &mut [u8]) -> Result<usize, TooSmall> { + let needed = self.total_len(token); + let Some(dst) = out.get_mut(..needed) else { + return Err(TooSmall { needed }); + }; + let sep = if token == Token::Cookie || token == Token::SetCookie { + b';' + } else if token == Token::UriArgs { + b'&' + } else { + b',' + }; + let mut at = 0usize; + for (i, f) in self.fragments(token).enumerate() { + if i > 0 { + if let Some(b) = dst.get_mut(at) { + *b = sep; + } + at = at.saturating_add(1); + } + let end = at.saturating_add(f.len()); + if let Some(d) = dst.get_mut(at..end) { + d.copy_from_slice(f); + } + at = end; + } + Ok(needed) + } + + /// The headers lws does not know, as (name, value), in the order their + /// names ended. A name is lowercase and ends with its `:`, as C keeps + /// it. + /// + /// ``` + /// use npro_h1::head::{Config, Head, Side}; + /// + /// let mut h = Head::new([0u8; 512], Side::Server, Config::default())?; + /// h.rx(b"GET / HTTP/1.1\r\nX-Foo: bar\r\n\r\n")?; + /// let u: Vec<(&[u8], &[u8])> = h.table().unknown_headers().collect(); + /// assert_eq!(u, [(&b"x-foo:"[..], &b"bar"[..])]); + /// # Ok::<(), Box<dyn core::error::Error>>(()) + /// ``` + #[must_use] + pub fn unknown_headers(&self) -> UnknownHeaders<'_> { + UnknownHeaders { + data: self.data.as_ref(), + next: match self.unknowns { + Unknowns::Linked { head, .. } => Some(head), + Unknowns::None => None, + }, + } + } + + /// The value of the unknown header `name`, given lowercase with its + /// `:`: C's `lws_hdr_custom_copy()`. + #[must_use] + pub fn unknown(&self, name: &[u8]) -> Option<&[u8]> { + self.unknown_headers() + .find(|(n, _)| *n == name) + .map(|(_, v)| v) + } + + /// Adds `value` to `token`, as another piece if it is there already: + /// C's `lws_hdr_simple_create()`, how a client keeps its own request. + /// An empty `value` takes the token out of the table. + /// + /// Unlike C, which can leave part of a value behind when it runs out + /// of room, it adds all of it or nothing. + /// + /// # Errors + /// + /// [`Full`], changing nothing, if there is no room for it. + pub fn create(&mut self, token: Token, value: &[u8]) -> Result<(), Full> { + if value.is_empty() { + if let Some(s) = self.index.get_mut(token.index()) { + *s = Slot::Absent; + } + return Ok(()); + } + // the value and its NUL + let len = u16::try_from(value.len()).map_err(|_| Full)?; + let end = self + .pos + .checked_add(len) + .and_then(|e| e.checked_add(1)) + .ok_or(Full)?; + if end > self.capacity { + return Err(Full); + } + self.start_fragment(token)?; + let start = self.pos; + let dst = self + .data + .as_mut() + .get_mut(usize::from(start)..usize::from(end)) + .ok_or(Full)?; + let (v, nul) = dst.split_at_mut(value.len()); + v.copy_from_slice(value); + nul.fill(0); + self.pos = end; + if let Some(f) = self.current_mut() { + f.len = len; + } + Ok(()) + } + + /// Marks where a client's response starts, after its own request: C's + /// `lws_header_table_rx_snapshot()`. [`rewind`](Self::rewind) goes + /// back to here. + pub const fn snapshot(&mut self) { + self.mark = Mark { + pos: self.pos, + nfrag: self.nfrag, + unknowns: self.unknowns, + }; + } + + /// Drops everything added since the [`snapshot`](Self::snapshot), as a + /// client does with an interim (1xx) response: C's + /// `lws_header_table_rx_rewind()`. + pub fn rewind(&mut self) { + let mark = self.mark; + for s in &mut self.index { + if let Slot::Present(f) = *s { + if f.0 > mark.nfrag { + *s = Slot::Absent; + } + } + } + for (n, f) in self.frags.iter_mut().enumerate() { + if n <= usize::from(mark.nfrag) { + if f.next.is_some_and(|x| x.0 > mark.nfrag) { + f.next = None; + } + } else { + *f = Frag::default(); + } + } + self.nfrag = mark.nfrag; + self.pos = mark.pos; + self.unknowns = mark.unknowns; + // the restored tail's link still names a dropped record + if let Unknowns::Linked { tail, .. } = self.unknowns { + self.write_be(tail.saturating_add(4), &[0; 4]); + } + } + + // What the head parser builds the table with. + + fn slot(&self, token: Token) -> Slot { + self.index + .get(token.index()) + .copied() + .unwrap_or(Slot::Absent) + } + + pub(crate) const fn pos(&self) -> u16 { + self.pos + } + + pub(crate) const fn has_fragments(&self) -> bool { + self.nfrag != 0 + } + + /// Whether another byte fits: C's `lws_pos_in_bounds()`. + pub(crate) const fn has_room(&self) -> bool { + self.pos < self.capacity + } + + /// The byte at `at`, if it is in the table. + pub(crate) fn byte(&self, at: u16) -> Option<u8> { + self.data.as_ref().get(usize::from(at)).copied() + } + + /// Lays down `c` past everything, outside any fragment. + pub(crate) fn push(&mut self, c: u8) -> Result<(), Full> { + if !self.has_room() { + return Err(Full); + } + let slot = self + .data + .as_mut() + .get_mut(usize::from(self.pos)) + .ok_or(Full)?; + *slot = c; + self.pos = self.pos.checked_add(1).ok_or(Full)?; + Ok(()) + } + + /// Sets the end of the table back to `pos`, dropping what is past it. + pub(crate) fn truncate(&mut self, pos: u16) { + self.pos = self.pos.min(pos); + } + + fn current_mut(&mut self) -> Option<&mut Frag> { + self.frags.get_mut(usize::from(self.nfrag)) + } + + /// The fragment being filled. + pub(crate) fn current(&self) -> &[u8] { + let Some(f) = self.frags.get(usize::from(self.nfrag)) else { + return &[]; + }; + let start = usize::from(f.offset); + self.data + .as_ref() + .get(start..start.saturating_add(usize::from(f.len))) + .unwrap_or_default() + } + + /// How long the fragment being filled is. + pub(crate) fn current_len(&self) -> u16 { + self.frags.get(usize::from(self.nfrag)).map_or(0, |f| f.len) + } + + /// Whether the fragment being filled is `token`'s first. + pub(crate) fn filling_first(&self, token: Token) -> bool { + self.slot(token) == Slot::Present(FragId(self.nfrag)) + } + + /// How long `token`'s first fragment is. + pub(crate) fn first_len(&self, token: Token) -> u16 { + match self.slot(token) { + Slot::Present(f) => self.frags.get(f.slot()).map_or(0, |f| f.len), + Slot::Absent => 0, + } + } + + /// Adds `c` to the fragment being filled, `max` being the most it may + /// hold: C's `issue_char()`. A NUL is the fragment's end, and does not + /// count against `max`. + pub(crate) fn issue(&mut self, c: u8, max: Option<u16>) -> Result<(), Full> { + if !self.has_room() { + return Err(Full); + } + let len = self.current_len(); + if c != 0 && max.is_some_and(|m| len >= m) { + return Err(Full); + } + self.push(c)?; + if let Some(f) = self.current_mut() { + f.len = f.len.saturating_add(1); + } + Ok(()) + } + + /// Takes the last byte of the fragment being filled out of its length, + /// leaving it in the table: how C leaves the NUL after a value. + pub(crate) fn uncount(&mut self) { + if let Some(f) = self.current_mut() { + f.len = f.len.saturating_sub(1); + } + } + + /// Begins a fragment for `token`, chained on to its last if it is + /// there already. Returns whether it was. + pub(crate) fn start_fragment(&mut self, token: Token) -> Result<bool, Full> { + let id = self.next_frag()?; + let pos = self.pos; + if let Some(f) = self.frags.get_mut(id.slot()) { + *f = Frag { + offset: pos, + len: 0, + next: None, + }; + } + self.nfrag = id.0; + let Some(slot) = self.index.get_mut(token.index()) else { + return Err(Full); + }; + let Slot::Present(mut last) = *slot else { + *slot = Slot::Present(id); + return Ok(false); + }; + // to the end of the chain, which only goes forward + while let Some(n) = self.frags.get(last.slot()).and_then(|f| f.next) { + if n <= last { + break; + } + last = n; + } + if let Some(f) = self.frags.get_mut(last.slot()) { + f.next = Some(id); + } + Ok(true) + } + + /// The next fragment slot, if there is one: C's check before moving + /// `nfrag` on. + fn next_frag(&self) -> Result<FragId, Full> { + let n = self.nfrag.checked_add(1).ok_or(Full)?; + if usize::from(n) >= Token::COUNT { + return Err(Full); + } + Ok(FragId(n)) + } + + /// Ends the urlarg being filled and begins the next, at `&` or `;`. + /// The caller has laid down the NUL that ends it. + pub(crate) fn next_arg(&mut self) -> Result<(), Full> { + let id = self.next_frag()?; + if !self.has_room() { + return Err(Full); + } + if let Some(f) = self.current_mut() { + f.next = Some(id); + } + self.nfrag = id.0; + let pos = self.pos; + if let Some(f) = self.current_mut() { + *f = Frag { + offset: pos, + len: 0, + next: None, + }; + } + Ok(()) + } + + /// Begins the urlargs, at the request target's `?`. The caller has + /// laid down the NUL that ends the path. The urlargs start a byte + /// past it, as in C. + pub(crate) fn start_args(&mut self) -> Result<(), Full> { + let id = self.next_frag()?; + let at = self.pos.checked_add(1).ok_or(Full)?; + if at >= self.capacity { + return Err(Full); + } + self.nfrag = id.0; + self.pos = at; + if let Some(f) = self.current_mut() { + *f = Frag { + offset: at, + len: 0, + next: None, + }; + } + if let Some(s) = self.index.get_mut(Token::UriArgs.index()) { + *s = Slot::Present(id); + } + Ok(()) + } + + /// Takes the path being filled back to the `/` before its last + /// segment, for a `/..`: the loop C has at each `/../`. + pub(crate) fn back_up_a_segment(&mut self) { + if self.current_len() <= 2 { + return; + } + self.pos = self.pos.saturating_sub(1); + self.uncount(); + loop { + self.pos = self.pos.saturating_sub(1); + self.uncount(); + if self.current_len() <= 1 || self.byte(self.pos) == Some(b'/') { + break; + } + } + } + + /// Begins an unknown header's record at the end of the table, its + /// eight bytes zeroed while there is room for them, as C does. + pub(crate) fn begin_record(&mut self) -> u16 { + let at = self.pos; + for _ in 0..RECORD { + if self.push(0).is_err() { + break; + } + } + at + } + + /// The name of the record at `rec` so far, everything past its eight + /// bytes. + pub(crate) fn record_name(&self, rec: u16) -> &[u8] { + let start = usize::from(rec.saturating_add(RECORD)); + self.data + .as_ref() + .get(start..usize::from(self.pos)) + .unwrap_or_default() + } + + /// The name of the record at `rec` has ended, here: its length goes in + /// its record, and the record on the end of the list. + pub(crate) fn end_record_name(&mut self, rec: u16) { + self.unknowns = match self.unknowns { + Unknowns::None => Unknowns::Linked { + head: rec, + tail: rec, + }, + Unknowns::Linked { head, tail } => { + self.write_be(tail.saturating_add(4), &u32::from(rec).to_be_bytes()); + Unknowns::Linked { head, tail: rec } + } + }; + let nlen = self.pos.saturating_sub(rec.saturating_add(RECORD)); + self.write_be(rec, &nlen.to_be_bytes()); + } + + /// The value of the record at `rec`, which started at `value`, has + /// ended, here. + pub(crate) fn end_record_value(&mut self, rec: u16, value: u16) { + let vlen = self.pos.saturating_sub(value); + self.write_be(rec.saturating_add(2), &vlen.to_be_bytes()); + } + + fn write_be(&mut self, at: u16, bytes: &[u8]) { + let at = usize::from(at); + if let Some(d) = self + .data + .as_mut() + .get_mut(at..at.saturating_add(bytes.len())) + { + d.copy_from_slice(bytes); + } + } +} + +/// The pieces of a token's value: see [`HeaderTable::fragments`]. +#[derive(Clone, Debug)] +pub struct Fragments<'a> { + data: &'a [u8], + frags: &'a [Frag; Token::COUNT], + next: Option<FragId>, +} + +impl<'a> Iterator for Fragments<'a> { + type Item = &'a [u8]; + + fn next(&mut self) -> Option<&'a [u8]> { + let id = self.next?; + let f = self.frags.get(id.slot())?; + // a chain only goes forward, as C's lws_ah_frag_next() insists + self.next = f.next.filter(|n| *n > id); + let start = usize::from(f.offset); + self.data + .get(start..start.saturating_add(usize::from(f.len))) + } +} + +/// The unknown headers: see [`HeaderTable::unknown_headers`]. +#[derive(Clone, Debug)] +pub struct UnknownHeaders<'a> { + data: &'a [u8], + next: Option<u16>, +} + +impl<'a> Iterator for UnknownHeaders<'a> { + type Item = (&'a [u8], &'a [u8]); + + fn next(&mut self) -> Option<Self::Item> { + let rec = usize::from(self.next?); + let head = self.data.get(rec..rec.saturating_add(8))?; + let (nlen, rest) = head.split_first_chunk::<2>()?; + let (vlen, link) = rest.split_first_chunk::<2>()?; + let link = u32::from_be_bytes(*link.first_chunk::<4>()?); + // the next record is further on, or there is none: never round + self.next = u16::try_from(link).ok().filter(|n| usize::from(*n) > rec); + let name = rec.saturating_add(8); + let value = name.saturating_add(usize::from(u16::from_be_bytes(*nlen))); + let end = value.saturating_add(usize::from(u16::from_be_bytes(*vlen))); + Some((self.data.get(name..value)?, self.data.get(value..end)?)) + } +} + +#[cfg(test)] +mod tests { + extern crate alloc; + + use super::*; + use alloc::vec; + use alloc::vec::Vec; + + #[test] + fn storage_past_the_cap_is_refused() { + assert!(HeaderTable::new(vec![0u8; MAX_CAPACITY]).is_ok()); + assert_eq!( + HeaderTable::new(vec![0u8; MAX_CAPACITY + 1]).err(), + Some(CapacityTooLarge) + ); + } + + #[test] + fn create_chains_and_copy_joins() { + let mut t = HeaderTable::new([0u8; 64]).unwrap(); + t.create(Token::Accept, b"a").unwrap(); + t.create(Token::Accept, b"bc").unwrap(); + t.create(Token::Cookie, b"x=1").unwrap(); + t.create(Token::Cookie, b"y=2").unwrap(); + assert_eq!(t.total_len(Token::Accept), 4); + let mut out = [0u8; 8]; + assert_eq!(t.copy(Token::Accept, &mut out), Ok(4)); + assert_eq!(&out[..4], b"a,bc"); + assert_eq!(t.copy(Token::Cookie, &mut out), Ok(7)); + assert_eq!(&out[..7], b"x=1;y=2"); + assert_eq!( + t.copy(Token::Cookie, &mut [0u8; 6]), + Err(TooSmall { needed: 7 }) + ); + // each value and its NUL + assert_eq!(t.used(), 2 + 3 + 4 + 4); + } + + #[test] + fn create_is_all_or_nothing() { + let mut t = HeaderTable::new([0u8; 8]).unwrap(); + t.create(Token::Host, b"abc").unwrap(); + assert_eq!(t.create(Token::Host, b"defg"), Err(Full)); + assert_eq!(t.first(Token::Host), Some(&b"abc"[..])); + assert_eq!(t.fragments(Token::Host).count(), 1); + t.create(Token::Host, b"").unwrap(); + assert!(!t.is_present(Token::Host)); + } + + #[test] + fn rewind_keeps_what_was_before_the_snapshot() { + let mut t = HeaderTable::new([0u8; 64]).unwrap(); + t.create(Token::ClientUri, b"/x").unwrap(); + t.create(Token::Accept, b"a").unwrap(); + t.snapshot(); + let used = t.used(); + t.create(Token::Accept, b"b").unwrap(); + t.create(Token::Host, b"h").unwrap(); + t.rewind(); + assert_eq!(t.used(), used); + assert!(!t.is_present(Token::Host)); + assert_eq!(t.fragments(Token::Accept).count(), 1); + t.create(Token::Accept, b"c").unwrap(); + let v: Vec<&[u8]> = t.fragments(Token::Accept).collect(); + assert_eq!(v, [&b"a"[..], b"c"]); + } +} diff --git a/crates/npro-h1/src/token.rs b/crates/npro-h1/src/token.rs new file mode 100644 index 0000000..ec77172 --- /dev/null +++ b/crates/npro-h1/src/token.rs @@ -0,0 +1,569 @@ +//! The headers lws knows by name: C's `enum lws_token_indexes`. +//! +//! C matches a name against its generated lextable, a trie over the +//! spellings in `lextable-strings.h`, one byte at a time. The trie is how +//! C finds them, not what they are: the spellings, their delimiters (the +//! `:` of a field name, the SP after a method, nothing after an h2 +//! pseudo-header) and the token each one stands for are the behaviour, and +//! those are here. [`lookup`] answers what the trie answers after each +//! byte of a name: matched, still a prefix of something, or nothing. +//! +//! The set is C's with every header option on, as C's default build has +//! it (`LWS_WITH_HTTP_UNCOMMON_HEADERS`, `LWS_ROLE_WS`, `LWS_ROLE_H2`), and +//! the indices are C's, so a table dumped from either side lines up. + +/// A header, or a piece of a head, that lws stores by token. +/// +/// Each variant names its C token. Some hold what is not a field's value: +/// the request target of each method, the request's version or the +/// response's status, and the urlargs. The last nine are C's +/// `_WSI_TOKEN_CLIENT_*`, where a client keeps its own request: they are +/// never matched, only created. +/// +/// ``` +/// use npro_h1::token::Token; +/// +/// assert_eq!(Token::Host.spelling(), b"host:"); +/// assert_eq!(Token::GetUri.spelling(), b"get "); +/// assert_eq!(Token::ALL.len(), Token::COUNT); +/// assert_eq!(Token::ALL[27], Token::ContentLength); +/// ``` +#[derive(Clone, Copy, Debug, PartialEq, Eq, PartialOrd, Ord, Hash)] +#[repr(u8)] +pub enum Token { + /// `WSI_TOKEN_GET_URI`: a GET's request target. + GetUri = 0, + /// `WSI_TOKEN_POST_URI`: a POST's request target. + PostUri, + /// `WSI_TOKEN_OPTIONS_URI`: an OPTIONS's request target. + OptionsUri, + /// `WSI_TOKEN_HOST`. + Host, + /// `WSI_TOKEN_CONNECTION`. + Connection, + /// `WSI_TOKEN_UPGRADE`. + Upgrade, + /// `WSI_TOKEN_ORIGIN`, and `Sec-WebSocket-Origin`, which C stores here. + Origin, + /// `WSI_TOKEN_DRAFT`: `Sec-WebSocket-Draft`. + WsDraft, + /// `WSI_TOKEN_CHALLENGE`: the empty line that ends a head, which the + /// lextable matches as the name `"\r\n"`. Never stored. + Challenge, + /// `WSI_TOKEN_EXTENSIONS`: `Sec-WebSocket-Extensions`. + WsExtensions, + /// `WSI_TOKEN_KEY1`: `Sec-WebSocket-Key1`, of the hixie drafts. + WsKey1, + /// `WSI_TOKEN_KEY2`: `Sec-WebSocket-Key2`, of the hixie drafts. + WsKey2, + /// `WSI_TOKEN_PROTOCOL`: `Sec-WebSocket-Protocol`. + WsProtocol, + /// `WSI_TOKEN_ACCEPT`: `Sec-WebSocket-Accept`. + WsAccept, + /// `WSI_TOKEN_NONCE`: `Sec-WebSocket-Nonce`. + WsNonce, + /// `WSI_TOKEN_HTTP`: on a server, the request line's version; on a + /// client, what follows a status line's `HTTP/1.1 `. + Http, + /// `WSI_TOKEN_HTTP2_SETTINGS`. + Http2Settings, + /// `WSI_TOKEN_HTTP_ACCEPT`. + Accept, + /// `WSI_TOKEN_HTTP_AC_REQUEST_HEADERS`: + /// `Access-Control-Request-Headers`. + AccessControlRequestHeaders, + /// `WSI_TOKEN_HTTP_IF_MODIFIED_SINCE`. + IfModifiedSince, + /// `WSI_TOKEN_HTTP_IF_NONE_MATCH`. + IfNoneMatch, + /// `WSI_TOKEN_HTTP_ACCEPT_ENCODING`. + AcceptEncoding, + /// `WSI_TOKEN_HTTP_ACCEPT_LANGUAGE`. + AcceptLanguage, + /// `WSI_TOKEN_HTTP_PRAGMA`. + Pragma, + /// `WSI_TOKEN_HTTP_CACHE_CONTROL`. + CacheControl, + /// `WSI_TOKEN_HTTP_AUTHORIZATION`. + Authorization, + /// `WSI_TOKEN_HTTP_COOKIE`. + Cookie, + /// `WSI_TOKEN_HTTP_CONTENT_LENGTH`. + ContentLength, + /// `WSI_TOKEN_HTTP_CONTENT_TYPE`. + ContentType, + /// `WSI_TOKEN_HTTP_DATE`. + Date, + /// `WSI_TOKEN_HTTP_RANGE`. + Range, + /// `WSI_TOKEN_HTTP_REFERER`. + Referer, + /// `WSI_TOKEN_KEY`: `Sec-WebSocket-Key`. + WsKey, + /// `WSI_TOKEN_VERSION`: `Sec-WebSocket-Version`. + WsVersion, + /// `WSI_TOKEN_SWORIGIN`: `Sec-WebSocket-Origin`. Matched, then stored + /// as [`Token::Origin`], as C does. + WsOrigin, + /// `WSI_TOKEN_HTTP_COLON_AUTHORITY`: h2 and h3's `:authority`. + ColonAuthority, + /// `WSI_TOKEN_HTTP_COLON_METHOD`. + ColonMethod, + /// `WSI_TOKEN_HTTP_COLON_PATH`. + ColonPath, + /// `WSI_TOKEN_HTTP_COLON_SCHEME`. + ColonScheme, + /// `WSI_TOKEN_HTTP_COLON_STATUS`. + ColonStatus, + /// `WSI_TOKEN_HTTP_ACCEPT_CHARSET`. + AcceptCharset, + /// `WSI_TOKEN_HTTP_ACCEPT_RANGES`. + AcceptRanges, + /// `WSI_TOKEN_HTTP_ACCESS_CONTROL_ALLOW_ORIGIN`. + AccessControlAllowOrigin, + /// `WSI_TOKEN_HTTP_AGE`. + Age, + /// `WSI_TOKEN_HTTP_ALLOW`. + Allow, + /// `WSI_TOKEN_HTTP_CONTENT_DISPOSITION`. + ContentDisposition, + /// `WSI_TOKEN_HTTP_CONTENT_ENCODING`. + ContentEncoding, + /// `WSI_TOKEN_HTTP_CONTENT_LANGUAGE`. + ContentLanguage, + /// `WSI_TOKEN_HTTP_CONTENT_LOCATION`. + ContentLocation, + /// `WSI_TOKEN_HTTP_CONTENT_RANGE`. + ContentRange, + /// `WSI_TOKEN_HTTP_ETAG`. + Etag, + /// `WSI_TOKEN_HTTP_EXPECT`. + Expect, + /// `WSI_TOKEN_HTTP_EXPIRES`. + Expires, + /// `WSI_TOKEN_HTTP_FROM`. + From, + /// `WSI_TOKEN_HTTP_IF_MATCH`. + IfMatch, + /// `WSI_TOKEN_HTTP_IF_RANGE`. + IfRange, + /// `WSI_TOKEN_HTTP_IF_UNMODIFIED_SINCE`. + IfUnmodifiedSince, + /// `WSI_TOKEN_HTTP_LAST_MODIFIED`. + LastModified, + /// `WSI_TOKEN_HTTP_LINK`. + Link, + /// `WSI_TOKEN_HTTP_LOCATION`. + Location, + /// `WSI_TOKEN_HTTP_MAX_FORWARDS`. + MaxForwards, + /// `WSI_TOKEN_HTTP_PROXY_AUTHENTICATE`. + ProxyAuthenticate, + /// `WSI_TOKEN_HTTP_PROXY_AUTHORIZATION`. + ProxyAuthorization, + /// `WSI_TOKEN_HTTP_REFRESH`. + Refresh, + /// `WSI_TOKEN_HTTP_RETRY_AFTER`. + RetryAfter, + /// `WSI_TOKEN_HTTP_SERVER`. + Server, + /// `WSI_TOKEN_HTTP_SET_COOKIE`. + SetCookie, + /// `WSI_TOKEN_HTTP_STRICT_TRANSPORT_SECURITY`. + StrictTransportSecurity, + /// `WSI_TOKEN_HTTP_TRANSFER_ENCODING`. + TransferEncoding, + /// `WSI_TOKEN_HTTP_USER_AGENT`. + UserAgent, + /// `WSI_TOKEN_HTTP_VARY`. + Vary, + /// `WSI_TOKEN_HTTP_VIA`. + Via, + /// `WSI_TOKEN_HTTP_WWW_AUTHENTICATE`. + WwwAuthenticate, + /// `WSI_TOKEN_PATCH_URI`: a PATCH's request target. + PatchUri, + /// `WSI_TOKEN_PUT_URI`: a PUT's request target. + PutUri, + /// `WSI_TOKEN_DELETE_URI`: a DELETE's request target. + DeleteUri, + /// `WSI_TOKEN_HTTP_URI_ARGS`: the request target's urlargs, the part + /// after its `?`, one fragment per `&` or `;` separated argument. + UriArgs, + /// `WSI_TOKEN_PROXY`. + Proxy, + /// `WSI_TOKEN_HTTP_X_REAL_IP`. + XRealIp, + /// `WSI_TOKEN_HTTP1_0`: on a client, what follows a status line's + /// `HTTP/1.0 `. + Http10, + /// `WSI_TOKEN_X_FORWARDED_FOR`. + XForwardedFor, + /// `WSI_TOKEN_CONNECT`: a CONNECT's request target. + Connect, + /// `WSI_TOKEN_HEAD_URI`: a HEAD's request target. + HeadUri, + /// `WSI_TOKEN_TE`. + Te, + /// `WSI_TOKEN_REPLAY_NONCE`: ACME's `Replay-Nonce`. + ReplayNonce, + /// `WSI_TOKEN_COLON_PROTOCOL`: RFC 8441's `:protocol`. + ColonProtocol, + /// `WSI_TOKEN_X_AUTH_TOKEN`. + XAuthToken, + /// `WSI_TOKEN_DSS_SIGNATURE`: `X-Amzn-Dss-Signature`. + DssSignature, + /// `_WSI_TOKEN_CLIENT_SENT_PROTOCOLS`: the ws subprotocols a client + /// asked for. + ClientSentProtocols, + /// `_WSI_TOKEN_CLIENT_PEER_ADDRESS`: the address a client connects to. + ClientPeerAddress, + /// `_WSI_TOKEN_CLIENT_URI`: the path a client asks for. + ClientUri, + /// `_WSI_TOKEN_CLIENT_HOST`: the `Host` a client sends. + ClientHost, + /// `_WSI_TOKEN_CLIENT_ORIGIN`: the `Origin` a client sends. + ClientOrigin, + /// `_WSI_TOKEN_CLIENT_METHOD`: a client's method. + ClientMethod, + /// `_WSI_TOKEN_CLIENT_IFACE`: the interface a client binds to. + ClientIface, + /// `_WSI_TOKEN_CLIENT_LOCALPORT`: the port a client binds to. + ClientLocalport, + /// `_WSI_TOKEN_CLIENT_ALPN`: the alpn a client offers. + ClientAlpn, +} + +impl Token { + /// How many tokens there are: C's `WSI_TOKEN_COUNT`. + pub const COUNT: usize = 97; + + /// Every token, in C's order, so `ALL[t as usize] == t`. + pub const ALL: [Self; Self::COUNT] = [ + Self::GetUri, + Self::PostUri, + Self::OptionsUri, + Self::Host, + Self::Connection, + Self::Upgrade, + Self::Origin, + Self::WsDraft, + Self::Challenge, + Self::WsExtensions, + Self::WsKey1, + Self::WsKey2, + Self::WsProtocol, + Self::WsAccept, + Self::WsNonce, + Self::Http, + Self::Http2Settings, + Self::Accept, + Self::AccessControlRequestHeaders, + Self::IfModifiedSince, + Self::IfNoneMatch, + Self::AcceptEncoding, + Self::AcceptLanguage, + Self::Pragma, + Self::CacheControl, + Self::Authorization, + Self::Cookie, + Self::ContentLength, + Self::ContentType, + Self::Date, + Self::Range, + Self::Referer, + Self::WsKey, + Self::WsVersion, + Self::WsOrigin, + Self::ColonAuthority, + Self::ColonMethod, + Self::ColonPath, + Self::ColonScheme, + Self::ColonStatus, + Self::AcceptCharset, + Self::AcceptRanges, + Self::AccessControlAllowOrigin, + Self::Age, + Self::Allow, + Self::ContentDisposition, + Self::ContentEncoding, + Self::ContentLanguage, + Self::ContentLocation, + Self::ContentRange, + Self::Etag, + Self::Expect, + Self::Expires, + Self::From, + Self::IfMatch, + Self::IfRange, + Self::IfUnmodifiedSince, + Self::LastModified, + Self::Link, + Self::Location, + Self::MaxForwards, + Self::ProxyAuthenticate, + Self::ProxyAuthorization, + Self::Refresh, + Self::RetryAfter, + Self::Server, + Self::SetCookie, + Self::StrictTransportSecurity, + Self::TransferEncoding, + Self::UserAgent, + Self::Vary, + Self::Via, + Self::WwwAuthenticate, + Self::PatchUri, + Self::PutUri, + Self::DeleteUri, + Self::UriArgs, + Self::Proxy, + Self::XRealIp, + Self::Http10, + Self::XForwardedFor, + Self::Connect, + Self::HeadUri, + Self::Te, + Self::ReplayNonce, + Self::ColonProtocol, + Self::XAuthToken, + Self::DssSignature, + Self::ClientSentProtocols, + Self::ClientPeerAddress, + Self::ClientUri, + Self::ClientHost, + Self::ClientOrigin, + Self::ClientMethod, + Self::ClientIface, + Self::ClientLocalport, + Self::ClientAlpn, + ]; + + /// The request methods C knows, each the token holding its request + /// target: C's `methods[]`, in its order. + pub const METHODS: [Self; 8] = [ + Self::GetUri, + Self::PostUri, + Self::OptionsUri, + Self::PutUri, + Self::PatchUri, + Self::DeleteUri, + Self::Connect, + Self::HeadUri, + ]; + + /// C's index of the token. + #[must_use] + #[expect( + clippy::as_conversions, + reason = "a fieldless repr(u8) enum's discriminant, which is C's index, widened" + )] + pub const fn index(self) -> usize { + self as usize + } + + /// Whether the token holds a request method's target. + #[must_use] + pub fn is_method(self) -> bool { + Self::METHODS.contains(&self) + } + + /// How the token is spelled in C's lextable, lowercase, with its + /// delimiter: C's `lws_token_to_string()`. Empty for the tokens that + /// are never matched, the client's own. + #[must_use] + pub const fn spelling(self) -> &'static [u8] { + match self { + Self::GetUri => b"get ", + Self::PostUri => b"post ", + Self::OptionsUri => b"options ", + Self::Host => b"host:", + Self::Connection => b"connection:", + Self::Upgrade => b"upgrade:", + Self::Origin => b"origin:", + Self::WsDraft => b"sec-websocket-draft:", + Self::Challenge => b"\r\n", + Self::WsExtensions => b"sec-websocket-extensions:", + Self::WsKey1 => b"sec-websocket-key1:", + Self::WsKey2 => b"sec-websocket-key2:", + Self::WsProtocol => b"sec-websocket-protocol:", + Self::WsAccept => b"sec-websocket-accept:", + Self::WsNonce => b"sec-websocket-nonce:", + Self::Http => b"http/1.1 ", + Self::Http2Settings => b"http2-settings:", + Self::Accept => b"accept:", + Self::AccessControlRequestHeaders => b"access-control-request-headers:", + Self::IfModifiedSince => b"if-modified-since:", + Self::IfNoneMatch => b"if-none-match:", + Self::AcceptEncoding => b"accept-encoding:", + Self::AcceptLanguage => b"accept-language:", + Self::Pragma => b"pragma:", + Self::CacheControl => b"cache-control:", + Self::Authorization => b"authorization:", + Self::Cookie => b"cookie:", + Self::ContentLength => b"content-length:", + Self::ContentType => b"content-type:", + Self::Date => b"date:", + Self::Range => b"range:", + Self::Referer => b"referer:", + Self::WsKey => b"sec-websocket-key:", + Self::WsVersion => b"sec-websocket-version:", + Self::WsOrigin => b"sec-websocket-origin:", + Self::ColonAuthority => b":authority", + Self::ColonMethod => b":method", + Self::ColonPath => b":path", + Self::ColonScheme => b":scheme", + Self::ColonStatus => b":status", + Self::AcceptCharset => b"accept-charset:", + Self::AcceptRanges => b"accept-ranges:", + Self::AccessControlAllowOrigin => b"access-control-allow-origin:", + Self::Age => b"age:", + Self::Allow => b"allow:", + Self::ContentDisposition => b"content-disposition:", + Self::ContentEncoding => b"content-encoding:", + Self::ContentLanguage => b"content-language:", + Self::ContentLocation => b"content-location:", + Self::ContentRange => b"content-range:", + Self::Etag => b"etag:", + Self::Expect => b"expect:", + Self::Expires => b"expires:", + Self::From => b"from:", + Self::IfMatch => b"if-match:", + Self::IfRange => b"if-range:", + Self::IfUnmodifiedSince => b"if-unmodified-since:", + Self::LastModified => b"last-modified:", + Self::Link => b"link:", + Self::Location => b"location:", + Self::MaxForwards => b"max-forwards:", + Self::ProxyAuthenticate => b"proxy-authenticate:", + Self::ProxyAuthorization => b"proxy-authorization:", + Self::Refresh => b"refresh:", + Self::RetryAfter => b"retry-after:", + Self::Server => b"server:", + Self::SetCookie => b"set-cookie:", + Self::StrictTransportSecurity => b"strict-transport-security:", + Self::TransferEncoding => b"transfer-encoding:", + Self::UserAgent => b"user-agent:", + Self::Vary => b"vary:", + Self::Via => b"via:", + Self::WwwAuthenticate => b"www-authenticate:", + Self::PatchUri => b"patch", + Self::PutUri => b"put", + Self::DeleteUri => b"delete", + Self::UriArgs => b"uri-args", + Self::Proxy => b"proxy ", + Self::XRealIp => b"x-real-ip:", + Self::Http10 => b"http/1.0 ", + Self::XForwardedFor => b"x-forwarded-for:", + Self::Connect => b"connect ", + Self::HeadUri => b"head ", + Self::Te => b"te:", + Self::ReplayNonce => b"replay-nonce:", + Self::ColonProtocol => b":protocol", + Self::XAuthToken => b"x-auth-token:", + Self::DssSignature => b"x-amzn-dss-signature:", + Self::ClientSentProtocols + | Self::ClientPeerAddress + | Self::ClientUri + | Self::ClientHost + | Self::ClientOrigin + | Self::ClientMethod + | Self::ClientIface + | Self::ClientLocalport + | Self::ClientAlpn => b"", + } + } +} + +/// What C's lextable says of a name so far. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum Lookup { + /// The name is the whole spelling of this token. + Matched(Token), + /// The name is the start of at least one spelling, and the whole of + /// none. + Prefix, + /// The name is the start of no spelling. + Nothing, +} + +/// What C's lextable says of `name`, the lowercased bytes of a name so far. +/// +/// No spelling is the start of another, so a name matches as soon as it is +/// the whole of one, as the trie's terminal is reached on its last byte. +/// +/// ``` +/// use npro_h1::token::{lookup, Lookup, Token}; +/// +/// assert_eq!(lookup(b"hos"), Lookup::Prefix); +/// assert_eq!(lookup(b"host:"), Lookup::Matched(Token::Host)); +/// assert_eq!(lookup(b"host "), Lookup::Nothing); +/// assert_eq!(lookup(b"put"), Lookup::Matched(Token::PutUri)); +/// ``` +#[must_use] +pub fn lookup(name: &[u8]) -> Lookup { + let mut prefix = false; + for t in Token::ALL { + let s = t.spelling(); + if s.is_empty() || !s.starts_with(name) { + continue; + } + if s.len() == name.len() { + return Lookup::Matched(t); + } + prefix = true; + } + if prefix { + Lookup::Prefix + } else { + Lookup::Nothing + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn all_is_in_cs_order() { + for (i, t) in Token::ALL.iter().enumerate() { + assert_eq!(t.index(), i, "{t:?}"); + } + } + + /// The trie's terminal can only be reached on a spelling's last byte if + /// no spelling is the start of another: then a name is matched exactly + /// when it is a whole spelling, which is what `lookup()` relies on. + #[test] + fn no_spelling_is_the_start_of_another() { + for a in Token::ALL { + for b in Token::ALL { + let (sa, sb) = (a.spelling(), b.spelling()); + if a != b && !sa.is_empty() { + assert!(!sb.starts_with(sa), "{a:?} starts {b:?}"); + } + } + } + } + + #[test] + fn every_spelling_is_lowercase() { + for t in Token::ALL { + assert!(!t.spelling().iter().any(u8::is_ascii_uppercase), "{t:?}"); + } + } + + #[test] + fn every_spelling_matches_and_every_shorter_start_is_a_prefix() { + for t in Token::ALL { + let s = t.spelling(); + if s.is_empty() { + continue; + } + assert_eq!(lookup(s), Lookup::Matched(t)); + for n in 1..s.len() { + assert_eq!(lookup(&s[..n]), Lookup::Prefix); + } + } + } +} diff --git a/deny.toml b/deny.toml index d5f0f59..d19e6b0 100644 --- a/deny.toml +++ b/deny.toml @@ -38,6 +38,7 @@ allow = [ # the workspace "npro", "npro-core", + "npro-h1", "npro-fuzz", "npro-test", diff --git a/scripts/ci.sh b/scripts/ci.sh index a066593..6210e53 100755 --- a/scripts/ci.sh +++ b/scripts/ci.sh @@ -43,7 +43,7 @@ cargo "+$msrv" check --workspace --all-targets --all-features --locked echo "== no_std" # the sans-IO crates must build for a target with no std at all -for c in npro-core; do +for c in npro-core npro-h1; do cargo build -p "$c" --all-features --target thumbv7em-none-eabihf --locked done diff --git a/scripts/sai.sh b/scripts/sai.sh index ec873bd..a5d38b8 100755 --- a/scripts/sai.sh +++ b/scripts/sai.sh @@ -32,7 +32,7 @@ export PATH jobs="${SAI_PARALLEL:-4}" # the sans-IO crates, which must build with no std at all -nostd_crates="npro-core" +nostd_crates="npro-core npro-h1" nostd_targets="thumbv6m-none-eabi thumbv7em-none-eabihf riscv32imc-unknown-none-elf" . scripts/require.sh
Page fetched 0s ago, creation time: 7ms (vhost etag hits: 0%, cache hits: 0%)