Module: DS::Util::Strings
- Included in:
- DS::Util
- Defined in:
- lib/ds/util/strings.rb
Constant Summary collapse
- TERMINAL_PUNCT_REGEX =
TERMINAL_PUNCT_REGEX matches strings terminated by any of
.,;:?! %r{\s*([.,;:!]+)("?)$}- ELLIPSIS_REGEX =
ELLIPSIS_REGEX matches strings terminated by
... %r{\.\.\."?$}- ABBREV_REGEX =
ABBREV_REGEX matches values like 'N.T.', 'O.T.'
%r{\W[A-Z]\.$}- FINAL_QUESTION_REGEX =
Final ? regex
%r{\s*\?(\s*"?\s*)$}
Instance Method Summary collapse
-
#clean_string(string, terminator: nil, force: false) ⇒ String
This method calls.
- #clean_white_space(value) ⇒ Object
-
#convert_mets_superscript(value) ⇒ Object
converts encoded DS 1.0 encoded superscripts to parenthetical values; e.g., 'XVI#^4/4#' is converted to 'XVI(4/4)'.
-
#escape_pipes(value) ⇒ Object
Escape pipe characters in source strings so split operations can avoid splitting on them.
-
#fix_double_periods(value) ⇒ String
Replace any sequence of two '..' with a single period.
- #is_url?(value) ⇒ Boolean
-
#normalize_string(value) ⇒ String
Strip and replace all sequences of white space with single spaces and apply Unicode normalization.
- #remove_brackets(value) ⇒ Object
-
#terminate(str, terminator: '.', force: false) ⇒ String
Add termination to string if it lacks terminal punctuation.
-
#unicode_normalize(value, form = :nfc) ⇒ String
Return the string using unicode normalization form
form.
Instance Method Details
#clean_string(string, terminator: nil, force: false) ⇒ String
This method calls
convert_mets_superscriptremove_bracketsfix_double_periodsescape_pipesnormalize_string
If terminator is non-nil, the method removes any trailing
punctuation and whitespace and appends terminator.
Set terminator to `` (empty string) to remove trailing
punctuation.
26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 |
# File 'lib/ds/util/strings.rb', line 26 def clean_string string, terminator: nil, force: false normal = normalize_string( escape_pipes( fix_double_periods( remove_brackets( convert_mets_superscript(string.to_s) ) ) ) ) return normal if terminator.nil? cleaned = terminate normal, terminator: terminator, force: force # keep cleaning until no changes are made return clean_string cleaned unless cleaned == string cleaned end |
#clean_white_space(value) ⇒ Object
140 141 142 |
# File 'lib/ds/util/strings.rb', line 140 def clean_white_space value value.to_s.strip.gsub(%r{\s+}, ' ') end |
#convert_mets_superscript(value) ⇒ Object
converts encoded DS 1.0 encoded superscripts to parenthetical values; e.g., 'XVI#^4/4#' is converted to 'XVI(4/4)'
129 130 131 |
# File 'lib/ds/util/strings.rb', line 129 def convert_mets_superscript value value.to_s.gsub(%r{#\^([^#]+)#}, '(\1)') end |
#escape_pipes(value) ⇒ Object
Escape pipe characters in source strings so split operations can avoid splitting on them.
136 137 138 |
# File 'lib/ds/util/strings.rb', line 136 def escape_pipes value value.gsub('|', '\|') end |
#fix_double_periods(value) ⇒ String
Replace any sequence of two '..' with a single period. Ellipses, that is, sequences of three periods '...', are ignored.
fix_double_periods('....') # => "...."
fix_double_periods('.. ..') # => ". ."
fix_double_periods('... ..') # => "... ."
fix_double_periods('... a..') # => "... a."
fix_double_periods('a... a..') # => "a... a."
185 186 187 |
# File 'lib/ds/util/strings.rb', line 185 def fix_double_periods value value.to_s.gsub(%r{(?<!\.)\.\.(?!\.)}, '.') end |
#is_url?(value) ⇒ Boolean
189 190 191 |
# File 'lib/ds/util/strings.rb', line 189 def is_url? value value.to_s =~ URI::regexp end |
#normalize_string(value) ⇒ String
Strip and replace all sequences of white space with single spaces and apply Unicode normalization. NFC normalization is used for all strings except URLs, to which NFKC normalization is applied. See RFC 3987:
https://datatracker.ietf.org/doc/html/rfc3987#section-5.3.2.2
117 118 119 120 121 122 123 124 |
# File 'lib/ds/util/strings.rb', line 117 def normalize_string value form = is_url?(value) ? :nfkc : :nfc escape_pipes( clean_white_space( unicode_normalize(value, form) ) ) end |
#remove_brackets(value) ⇒ Object
168 169 170 |
# File 'lib/ds/util/strings.rb', line 168 def remove_brackets value value.to_s.strip.delete_prefix('[').delete_suffix(']') end |
#terminate(str, terminator: '.', force: false) ⇒ String
Add termination to string if it lacks terminal punctuation. Terminal punctuation is one of
. , ; : ? !
When :terminator is '' or nil, trailing punctuation isalways
removed.
Strings ending with ellipsis, '...' or '..."' are returned unaltered. This
behavior cannot be overridden with :force.
72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 |
# File 'lib/ds/util/strings.rb', line 72 def terminate str, terminator: '.', force: false str.strip! # DE 2022.08.12 Note the \s* to match and replace whitespace before # punctuation; this addresses a bug where some strings were returned # with trailing whitespace: 'value :' => 'value ' # TODO: Refactor? Two functions: strip_punctuation(), terminate() ?? # don't strip ellipses return str if str =~ ELLIPSIS_REGEX # don't strip final periods for strings like "N.T." return str if str =~ ABBREV_REGEX # don't strip final question marks return str if str =~ FINAL_QUESTION_REGEX # if :terminator is '' or nil, remove any terminal punctuation return str.sub TERMINAL_PUNCT_REGEX, '\2' if terminator.blank? # str is already terminated return str if str.end_with? terminator return str if str.end_with? %Q{#{terminator}"} # if string ends with '?', don't add terminator return str if str.end_with? '?' # str lacks terminal punctuation; add it; # \\1 => keep final '"' (double-quote) return str.sub %r{("?)$}, "#{terminator}\\1" if str !~ TERMINAL_PUNCT_REGEX # str has to have exact terminal punctuation # \\1 => keep final '"' (double-quote) return str.sub TERMINAL_PUNCT_REGEX, "#{terminator}\\2" if force # string has some terminal punctuation; return it str end |
#unicode_normalize(value, form = :nfc) ⇒ String
Return the string using unicode normalization form form.
Use NFC normalization by default. NFC normalization is
recommended best practice. See
https://www.honeybadger.io/blog/ruby-unicode-normalization/
In short: NFC should be used for most strings, but NFKC for URLs. See RFC 3987:
https://datatracker.ietf.org/doc/html/rfc3987#section-5.3.2.2
Wikibase uses NFC normalization:
164 165 166 |
# File 'lib/ds/util/strings.rb', line 164 def unicode_normalize value, form = :nfc value.to_s.unicode_normalize form end |