Module: TranslationDiff::Markup
- Defined in:
- lib/translation_diff/markup.rb
Overview
Entity references, and a < that opens no tag: the two places a document's text is not the text ox reports.
Defined Under Namespace
Classes: Resolver
Constant Summary collapse
- TAG_OPENER =
A
<opens a tag only when an element name, a closing name, a declaration or an instruction follows it. %r{[A-Za-z!?]|/[A-Za-z]}- AMBIGUOUS =
The two characters escaping has to move: a lone
<, and an&that would read as an escape this module wrote. /&(?=(?:amp;)*lt;)|<(?!#{TAG_OPENER})/- ESCAPED_ANGLE =
<was a lone<; every furtheramp;is a level the source itself wrote and escaping pushed up by one. /&((?:amp;)*)lt;/- DECODABLE =
Named and numeric alike, so nothing an
&opens survives to be escaped again and rendered as its own spelling. /&(?:[A-Za-z][A-Za-z0-9]*|#\d+|#[xX]\h+);/- NAMED =
Which of the two decoders an entity belongs to: Ox knows every HTML5 name, CGI knows both numeric forms.
/\A&([A-Za-z][A-Za-z0-9]*);\z/- ENCODED =
Only the two characters that are unsafe in HTML text; every other decoded character is left as the character it is.
{ "&" => "&", "<" => "<" }.freeze
- ENCODABLE =
/[&<]/- TRANSLATED_ENCODABLE =
An entity the round trip already produced must not be escaped a second time, and a
<shaped like a tag is trusted the same way a source tag already is -- everything else a provider sent back is untrusted new text. /&(?:amp|lt|gt);|&|<(?!#{TAG_OPENER})/
Class Method Summary collapse
-
.decode_entities(text) ⇒ Object
What a provider is sent is text, so it gets the characters; one left-to-right pass, so nothing is decoded twice.
-
.decoded(entity) ⇒ Object
An entity neither decoder knows stays as it arrived, and so does a surrogate: that decodes to invalid UTF-8.
-
.encode_entities(text) ⇒ Object
What a document renders is markup, so text that changed is made safe again -- and only where it is unsafe.
-
.encode_translation(text) ⇒ Object
What a translation renders as: unlike #encode_entities, this leaves a provider's own reproduced tags alone.
-
.escape_bare_angles(source) ⇒ Object
Hands back markup
oxcan parse: same document, with every lone<written as the entity it should have been. -
.named(name) ⇒ Object
A document repeats the same handful of names, and only a name that resolved is kept, so the table cannot be grown.
-
.resolve(name) ⇒ Object
One well-formed entity alone in an element is the only input Ox decodes safely -- prose with a lone
&raises. - .resolved ⇒ Object
-
.restore_bare_angles(rendered) ⇒ Object
The exact inverse: one
amp;off every escaped angle, and the angles with none left were the lone ones.
Class Method Details
.decode_entities(text) ⇒ Object
What a provider is sent is text, so it gets the characters; one left-to-right pass, so nothing is decoded twice.
41 |
# File 'lib/translation_diff/markup.rb', line 41 def self.decode_entities(text) = text.gsub(DECODABLE) { |entity| decoded(entity) } |
.decoded(entity) ⇒ Object
An entity neither decoder knows stays as it arrived, and so does a surrogate: that decodes to invalid UTF-8.
44 45 46 47 48 |
# File 'lib/translation_diff/markup.rb', line 44 def self.decoded(entity) name = entity[NAMED, 1] plain = name ? named(name) : CGI.unescapeHTML(entity) plain.valid_encoding? ? plain : entity end |
.encode_entities(text) ⇒ Object
What a document renders is markup, so text that changed is made safe again -- and only where it is unsafe.
68 |
# File 'lib/translation_diff/markup.rb', line 68 def self.encode_entities(text) = text.gsub(ENCODABLE, ENCODED) |
.encode_translation(text) ⇒ Object
What a translation renders as: unlike #encode_entities, this leaves a provider's own reproduced tags alone.
75 76 77 |
# File 'lib/translation_diff/markup.rb', line 75 def self.encode_translation(text) text.gsub(TRANSLATED_ENCODABLE) { |match| match.length == 1 ? ENCODED[match] : match } end |
.escape_bare_angles(source) ⇒ Object
Hands back markup ox can parse: same document, with every lone < written as the entity it should have been.
28 29 30 |
# File 'lib/translation_diff/markup.rb', line 28 def self.(source) source.gsub(AMBIGUOUS) { |ambiguous| ambiguous == "&" ? "&" : "<" } end |
.named(name) ⇒ Object
A document repeats the same handful of names, and only a name that resolved is kept, so the table cannot be grown.
51 |
# File 'lib/translation_diff/markup.rb', line 51 def self.named(name) = resolved[name] || resolve(name) |
.resolve(name) ⇒ Object
One well-formed entity alone in an element is the only input Ox decodes safely -- prose with a lone & raises.
56 57 58 59 60 61 62 63 64 65 |
# File 'lib/translation_diff/markup.rb', line 56 def self.resolve(name) entity = "&#{name};" resolver = Resolver.new Ox.sax_html(resolver, StringIO.new("<e>#{entity}</e>")) return entity if resolver.text.nil? || resolver.text == entity resolved[name] = resolver.text rescue StandardError entity end |
.resolved ⇒ Object
53 |
# File 'lib/translation_diff/markup.rb', line 53 def self.resolved = @resolved ||= {} |
.restore_bare_angles(rendered) ⇒ Object
The exact inverse: one amp; off every escaped angle, and the angles with none left were the lone ones.
33 34 35 36 37 38 |
# File 'lib/translation_diff/markup.rb', line 33 def self.(rendered) rendered.gsub(ESCAPED_ANGLE) do levels = Regexp.last_match(1) levels.empty? ? "<" : "&#{levels.delete_prefix('amp;')}lt;" end end |