Module: TranslationDiff::Markup

Defined in:
lib/translation_diff/markup.rb

Overview

Entity references, and a < that opens no tag: the two places a document's text is not the text ox reports.

Defined Under Namespace

Classes: Resolver

Constant Summary collapse

TAG_OPENER =

A < opens a tag only when an element name, a closing name, a declaration or an instruction follows it.

%r{[A-Za-z!?]|/[A-Za-z]}
AMBIGUOUS =

The two characters escaping has to move: a lone <, and an & that would read as an escape this module wrote.

/&(?=(?:amp;)*lt;)|<(?!#{TAG_OPENER})/
ESCAPED_ANGLE =

&lt; was a lone <; every further amp; is a level the source itself wrote and escaping pushed up by one.

/&((?:amp;)*)lt;/
DECODABLE =

Named and numeric alike, so nothing an & opens survives to be escaped again and rendered as its own spelling.

/&(?:[A-Za-z][A-Za-z0-9]*|#\d+|#[xX]\h+);/
NAMED =

Which of the two decoders an entity belongs to: Ox knows every HTML5 name, CGI knows both numeric forms.

/\A&([A-Za-z][A-Za-z0-9]*);\z/
ENCODED =

Only the two characters that are unsafe in HTML text; every other decoded character is left as the character it is.

{ "&" => "&amp;", "<" => "&lt;" }.freeze
ENCODABLE =
/[&<]/
TRANSLATED_ENCODABLE =

An entity the round trip already produced must not be escaped a second time, and a < shaped like a tag is trusted the same way a source tag already is -- everything else a provider sent back is untrusted new text.

/&(?:amp|lt|gt);|&|<(?!#{TAG_OPENER})/

Class Method Summary collapse

Class Method Details

.decode_entities(text) ⇒ Object

What a provider is sent is text, so it gets the characters; one left-to-right pass, so nothing is decoded twice.



41
# File 'lib/translation_diff/markup.rb', line 41

def self.decode_entities(text) = text.gsub(DECODABLE) { |entity| decoded(entity) }

.decoded(entity) ⇒ Object

An entity neither decoder knows stays as it arrived, and so does a surrogate: that decodes to invalid UTF-8.



44
45
46
47
48
# File 'lib/translation_diff/markup.rb', line 44

def self.decoded(entity)
  name = entity[NAMED, 1]
  plain = name ? named(name) : CGI.unescapeHTML(entity)
  plain.valid_encoding? ? plain : entity
end

.encode_entities(text) ⇒ Object

What a document renders is markup, so text that changed is made safe again -- and only where it is unsafe.



68
# File 'lib/translation_diff/markup.rb', line 68

def self.encode_entities(text) = text.gsub(ENCODABLE, ENCODED)

.encode_translation(text) ⇒ Object

What a translation renders as: unlike #encode_entities, this leaves a provider's own reproduced tags alone.



75
76
77
# File 'lib/translation_diff/markup.rb', line 75

def self.encode_translation(text)
  text.gsub(TRANSLATED_ENCODABLE) { |match| match.length == 1 ? ENCODED[match] : match }
end

.escape_bare_angles(source) ⇒ Object

Hands back markup ox can parse: same document, with every lone < written as the entity it should have been.



28
29
30
# File 'lib/translation_diff/markup.rb', line 28

def self.escape_bare_angles(source)
  source.gsub(AMBIGUOUS) { |ambiguous| ambiguous == "&" ? "&amp;" : "&lt;" }
end

.named(name) ⇒ Object

A document repeats the same handful of names, and only a name that resolved is kept, so the table cannot be grown.



51
# File 'lib/translation_diff/markup.rb', line 51

def self.named(name) = resolved[name] || resolve(name)

.resolve(name) ⇒ Object

One well-formed entity alone in an element is the only input Ox decodes safely -- prose with a lone & raises.



56
57
58
59
60
61
62
63
64
65
# File 'lib/translation_diff/markup.rb', line 56

def self.resolve(name)
  entity = "&#{name};"
  resolver = Resolver.new
  Ox.sax_html(resolver, StringIO.new("<e>#{entity}</e>"))
  return entity if resolver.text.nil? || resolver.text == entity

  resolved[name] = resolver.text
rescue StandardError
  entity
end

.resolved ⇒ Object



53
# File 'lib/translation_diff/markup.rb', line 53

def self.resolved = @resolved ||= {}

.restore_bare_angles(rendered) ⇒ Object

The exact inverse: one amp; off every escaped angle, and the angles with none left were the lone ones.



33
34
35
36
37
38
# File 'lib/translation_diff/markup.rb', line 33

def self.restore_bare_angles(rendered)
  rendered.gsub(ESCAPED_ANGLE) do
    levels = Regexp.last_match(1)
    levels.empty? ? "<" : "&#{levels.delete_prefix('amp;')}lt;"
  end
end