Class: Hpricot::Scanner

Inherits:
Object
  • Object
show all
Defined in:
lib/hpricot/scanner.rb

Overview

Liberal tokenizer for HTML/XML, modelled on ext/hpricot_scan/hpricot_scan.rl.

Rules of the port:

* Never raise on malformed input. Anything unrecognised becomes text.
* Every byte of the input appears in exactly one token's `raw`.
* Attribute values are stored undecoded.

Instance Method Summary collapse

Constructor Details

#initialize(source, xml: false, fixup_tags: false, xhtml_strict: false, html_void: false, void_elements: nil) ⇒ Scanner

Returns a new instance of Scanner.



20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
# File 'lib/hpricot/scanner.rb', line 20

def initialize(source, xml: false, fixup_tags: false, xhtml_strict: false, html_void: false,
               void_elements: nil)
  @xml = xml
  @fixup_tags = fixup_tags
  @xhtml_strict = xhtml_strict
  # Consumed by TreeBuilder; accepted here so Hpricot.scan can pass one
  # option hash to both.
  @html_void = html_void
  @void_elements = void_elements

  # Lowest position from which no '>' (respectively ']') remains in the
  # document. Both start unknown. See gt_ahead?.
  @no_gt_from = nil
  @no_rbracket_from = nil

  # Which encoding the token strings carry.
  #
  # A BINARY source carries no information about what its bytes mean --
  # it is what File.binread and Zip::File#read hand back -- so tag the
  # output Encoding.default_external, which is what the C scanner did for
  # every input. Callers rely on this: language_file_handler reads .docx
  # parts out of a zip (binary) and passes the result straight to
  # HTMLEntities#decode, which raises on an ASCII-8BIT string.
  #
  # A source that declares a real encoding is believed, which is where
  # this departs from the C scanner. That tagged every node
  # default_external regardless, so parsing an ISO-8859-1 document
  # produced UTF-8-labelled strings that failed valid_encoding?.
  @out_encoding =
    if source.encoding == Encoding::ASCII_8BIT
      Encoding.default_external
    else
      source.encoding
    end

  # What the regexes run against.
  #
  # Matching has to tolerate bytes that are invalid for the source's
  # declared encoding -- mislabelled Latin-1/CP-1252 HTML is common, and
  # the C scanner, a byte-oriented ragel machine with no concept of
  # encoding, never cared. An ordinary regex match on such a string raises
  # ArgumentError, so those get a binary view.
  #
  # A source that is already valid in an ASCII-compatible encoding needs no
  # copy: StringScanner#pos and #skip are byte-oriented whatever the
  # encoding, so the offset arithmetic here is unaffected, and `\s`, `\S`
  # and the character classes used below are all ASCII-only in Ruby. That
  # matters because the copy is the size of the document -- 1.5MB of
  # copying per parse, on top of a second full copy this used to make to
  # re-tag the source for output.
  @scan_src =
    if source.encoding.ascii_compatible? && source.valid_encoding?
      source
    else
      source.b
    end

  @ss = StringScanner.new(@scan_src)
end

Instance Method Details

#tokens ⇒ Object



80
81
82
83
84
85
86
# File 'lib/hpricot/scanner.rb', line 80

def tokens
  @tokens ||= begin
    out = []
    out << next_token until @ss.eos?
    out
  end
end