Class: Hpricot::Scanner
- Inherits:
-
Object
- Object
- Hpricot::Scanner
- Defined in:
- lib/hpricot/scanner.rb
Overview
Liberal tokenizer for HTML/XML, modelled on ext/hpricot_scan/hpricot_scan.rl.
Rules of the port:
* Never raise on malformed input. Anything unrecognised becomes text.
* Every byte of the input appears in exactly one token's `raw`.
* Attribute values are stored undecoded.
Instance Method Summary collapse
-
#initialize(source, xml: false, fixup_tags: false, xhtml_strict: false, html_void: false, void_elements: nil) ⇒ Scanner
constructor
A new instance of Scanner.
- #tokens ⇒ Object
Constructor Details
#initialize(source, xml: false, fixup_tags: false, xhtml_strict: false, html_void: false, void_elements: nil) ⇒ Scanner
Returns a new instance of Scanner.
20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 |
# File 'lib/hpricot/scanner.rb', line 20 def initialize(source, xml: false, fixup_tags: false, xhtml_strict: false, html_void: false, void_elements: nil) @xml = xml @fixup_tags = @xhtml_strict = xhtml_strict # Consumed by TreeBuilder; accepted here so Hpricot.scan can pass one # option hash to both. @html_void = html_void @void_elements = void_elements # Lowest position from which no '>' (respectively ']') remains in the # document. Both start unknown. See gt_ahead?. @no_gt_from = nil @no_rbracket_from = nil # Which encoding the token strings carry. # # A BINARY source carries no information about what its bytes mean -- # it is what File.binread and Zip::File#read hand back -- so tag the # output Encoding.default_external, which is what the C scanner did for # every input. Callers rely on this: language_file_handler reads .docx # parts out of a zip (binary) and passes the result straight to # HTMLEntities#decode, which raises on an ASCII-8BIT string. # # A source that declares a real encoding is believed, which is where # this departs from the C scanner. That tagged every node # default_external regardless, so parsing an ISO-8859-1 document # produced UTF-8-labelled strings that failed valid_encoding?. @out_encoding = if source.encoding == Encoding::ASCII_8BIT Encoding.default_external else source.encoding end # What the regexes run against. # # Matching has to tolerate bytes that are invalid for the source's # declared encoding -- mislabelled Latin-1/CP-1252 HTML is common, and # the C scanner, a byte-oriented ragel machine with no concept of # encoding, never cared. An ordinary regex match on such a string raises # ArgumentError, so those get a binary view. # # A source that is already valid in an ASCII-compatible encoding needs no # copy: StringScanner#pos and #skip are byte-oriented whatever the # encoding, so the offset arithmetic here is unaffected, and `\s`, `\S` # and the character classes used below are all ASCII-only in Ruby. That # matters because the copy is the size of the document -- 1.5MB of # copying per parse, on top of a second full copy this used to make to # re-tag the source for output. @scan_src = if source.encoding.ascii_compatible? && source.valid_encoding? source else source.b end @ss = StringScanner.new(@scan_src) end |
Instance Method Details
#tokens ⇒ Object
80 81 82 83 84 85 86 |
# File 'lib/hpricot/scanner.rb', line 80 def tokens @tokens ||= begin out = [] out << next_token until @ss.eos? out end end |