Module: MCPClient::SchemaValidator
- Extended by:
- Composition, Dialects, EcmaPatterns, Evaluation, InputRequirements, Instances, KeywordScan, Normalization, References, Scalars, Shapes, UriReferences
- Defined in:
- lib/mcp_client/schema_validator.rb,
lib/mcp_client/schema_validator/shapes.rb,
lib/mcp_client/schema_validator/scalars.rb,
lib/mcp_client/schema_validator/dialects.rb,
lib/mcp_client/schema_validator/instances.rb,
lib/mcp_client/schema_validator/evaluation.rb,
lib/mcp_client/schema_validator/references.rb,
lib/mcp_client/schema_validator/annotations.rb,
lib/mcp_client/schema_validator/composition.rb,
lib/mcp_client/schema_validator/keyword_scan.rb,
lib/mcp_client/schema_validator/ecma_patterns.rb,
lib/mcp_client/schema_validator/normalization.rb,
lib/mcp_client/schema_validator/uri_references.rb,
lib/mcp_client/schema_validator/input_requirements.rb
Overview
Self-contained JSON Schema validator used to check a tool call result's structuredContent against the tool's declared outputSchema (MCP server/tools spec: "Clients SHOULD validate structured results against this schema"; the default schema dialect is JSON Schema 2020-12 per basic "JSON Schema Usage").
Supported keywords:
- type (single value or array of values), enum, const
- properties, required (objects)
- items, prefixItems (2020-12) / tuple-form items (draft-07, 2019-09), minItems, maxItems (arrays)
- minLength, maxLength, pattern (an ECMA-262 regular expression, translated to the Ruby expression that means the same thing) (strings)
- minimum, maximum, exclusiveMinimum, exclusiveMaximum (numbers)
- allOf, anyOf, oneOf, not, if/then/else (composition)
- $ref to a location inside the same schema document (
#,#/$defs/x,#/definitions/x, any JSON pointer, or a plain-name fragment#namenaming an$anchor/ a draft-07$id: "#name"of the referencing schema's own resource — a subschema whose$idis a URI starts a new resource), with $defs (2019-09, 2020-12) and definitions, which the modern dialects keep as the deprecated spelling of $defs; under draft-07 a $ref replaces its siblings, under 2019-09 and 2020-12 it applies alongside them - boolean schemas (true / false), at the root or as subschemas
MCP 2026-07-28 rules honoured here:
-
a schema without
$schemais 2020-12; the dialects in SUPPORTED_DIALECTS are accepted and any other declared dialect is an error (not a permissive pass); the keyword grammar follows the dialect (DIALECT_KEYWORDS): a keyword the dialect does not define is ignored, neither shape-checked nor reported; -
$ref(and$dynamicRef/$recursiveRef) values that do not point inside the document (network URIs, relative documents, urn:, file:) are never dereferenced, and a schema carrying one is rejected rather than treated as permissive; -
resource bounds: schema nesting depth, total subschema count,
$refchain length, number of nodes visited, number of errors produced and a per-validation time budget. Hitting a bound aborts the validation with one error: an aborted validation never reads as a pass. -
multipleOf (numbers), uniqueItems, contains with minContains / maxContains, additionalItems (draft-07, 2019-09) (arrays)
-
minProperties, maxProperties, patternProperties, additionalProperties, propertyNames, dependentRequired / dependentSchemas (2019-09, 2020-12) and draft-07 dependencies (objects)
-
unevaluatedItems, unevaluatedProperties at a node that produces every annotation they read: with no in-place applicator beside them (and, for items, no
contains) they areadditionalItems/additionalPropertiesover what the node did not name
What is left out is exactly what cannot be decided here:
unevaluatedItems / unevaluatedProperties where an in-place applicator
(allOf, anyOf, oneOf, if, $ref, dependentSchemas — or a contains, whose
matches annotate the items they matched) contributes annotations
collected across a whole composition, the dynamic references
($dynamicRef, $recursiveRef), whose target depends on the dynamic scope a
validation was entered through, and the two keywords that only annotate
in the default dialect (format, which full validators assert only in
format-assertion mode, and contentSchema). Those are ignored rather than
misapplied, so validation is best-effort there — it may accept data a
full validator would reject, but it does not reject data that conforms to
the schema, and an unevaluated keyword is never read as a match for a
non-monotonic composition (not, oneOf, if). So that this gap is never
silent, unsupported_keywords reports which unapplied validation
keywords a schema uses; callers surface them as a warning.
Defined Under Namespace
Modules: Composition, Dialects, EcmaPatterns, Evaluation, InputRequirements, Instances, KeywordScan, Normalization, References, Scalars, Shapes, UriReferences Classes: Aborted, Context, Evaluated, TooLarge, Walk
Constant Summary collapse
- DEFAULT_DIALECT =
The default dialect (basic "JSON Schema Usage": "When a schema does not include a $schema field, it defaults to JSON Schema 2020-12").
'https://json-schema.org/draft/2020-12/schema'- DRAFT_2019_09 =
'https://json-schema.org/draft/2019-09/schema'- DRAFT_07 =
'http://json-schema.org/draft-07/schema'- SUPPORTED_DIALECTS =
Dialects this validator accepts. Any other declared dialect is reported as unsupported ("MUST handle unsupported dialects gracefully by returning an appropriate error indicating the dialect is not supported").
[DEFAULT_DIALECT, DRAFT_2019_09, DRAFT_07].freeze
- MAX_SCHEMA_DEPTH =
Resource bounds ("Composition-Keyword Resource Use": implementations SHOULD apply a maximum schema depth, a cap on the total number of subschemas, or a per-validation time budget). A schema comes from the remote server, so all of them apply.
64- MAX_SUBSCHEMAS =
2000- MAX_REF_DEPTH =
32- MAX_NODE_VISITS =
100_000- MAX_ERRORS =
100- MAX_VALUE_INSPECT =
64- MAX_NODE_DEPTH =
How deeply the walk may descend into the instance. The descent into a child value is the walk's only recursion, so this is what drives the interpreter's stack: the instance nests a level per array item or property, and a recursive
$reffollows it down without bound (the hop budget restarts at every value, so that a recursive schema can describe deep data at all). Counting the descent lets the deepest instance abort with one error like any other exhausted budget.Only a step into a child value — an array item or a property value — counts, and only such a step costs a frame. The schemas a node applies to one value (a
$refhop, an allOf/anyOf/oneOf/not/if branch) neither count nor recurse: a recursive schema composed through a few$defsmixins applies several of them per instance level, and both counting and recursing on those would scale with the mixin count, so data a peer can legitimately send — its nesting underJSON.parse's default limit of 100 — would abort. validate_node applies them iteratively (see Evaluation), bounded by MAX_REF_DEPTH and MAX_NODE_VISITS, which is what already stops a schema recursing on one value forever.The bound therefore sits well above the deepest instance a peer can send, and below the number of levels the smallest stack this library runs on (a transport's reader thread) carries — a figure that no longer depends on how a schema is composed, since what one level costs is now fixed. validate still catches SystemStackError, as a backstop for a stack smaller than any this has been measured on rather than for anything a schema can provoke.
256- UNSUPPORTED_KEYWORDS =
JSON Schema keywords that affect validation but that this validator may not evaluate: the dynamic references where they are dynamic (a target the dynamic scope a validation was entered through could re-bind, which this validator does not track), and the two that only annotate in the default dialect (
format, asserted by full validators in format-assertion mode, andcontentSchema). Where one of them is left unevaluated, validation is partial: data may pass here that a full validator would reject.A dynamic reference is not in this position: it is evaluated, against the dynamic scope the validation records as it enters each resource (References#dynamic_binding).
Every other standard keyword is evaluated —
unevaluatedItemsandunevaluatedPropertiesfrom the annotations Evaluation collects across a whole composition. A standard assertion left unevaluated is not a smaller report but a wrong verdict: it makes a composition branch undecided, and an undecided branch is accepted wherever the composition is monotonic (allOf,anyOf), so the instance passes a schema that rejects it. %w[contentSchema format].freeze
- ANNOTATION_KEYWORDS =
Unsupported keywords that are annotations, not assertions (
formatis annotation-only in the default 2020-12 vocabulary,contentSchemaonly annotates): their presence decides nothing about an instance. %w[format contentSchema].freeze
- DYNAMIC_REFERENCE_KEYWORDS =
The references whose target may depend on the dynamic scope.
%w[$dynamicRef $recursiveRef].freeze
- PATTERN_MATCH_TIMEOUT =
Wall-clock budget for a single validate call (pattern matching and the walk itself). Schemas come from the remote server, so an expensive expression or a huge composition must not be able to monopolize the calling thread.
The budget is for the whole operation, not per match: a per-match limit multiplies, since the server also controls how many strings it sends (N array items under one pathological items.pattern costs N x limit).
1.0- MIN_PATTERN_MATCH_TIMEOUT =
Floor for an individual match's timeout, so a nearly-exhausted budget still makes progress rather than failing every remaining pattern.
0.01- MAX_PATTERN_LENGTH =
The longest
pattern(orpatternPropertieskey) a schema may carry. A pattern is translated and compiled before it is matched, and both cost what the peer's text costs; the depth and subschema bounds say nothing about one string, so its length is bounded on its own. 10_000- SUBSCHEMA_KEYWORDS =
Keywords whose value is a single subschema to walk (or, for
items, an array of positional subschemas in draft-07 / 2019-09). %w[ items contains additionalProperties additionalItems propertyNames not if then else unevaluatedItems unevaluatedProperties contentSchema ].freeze
- SUBSCHEMA_MAP_KEYWORDS =
Keywords whose value is a map of name => subschema.
%w[properties patternProperties $defs definitions dependentSchemas dependencies].freeze
- SUBSCHEMA_ARRAY_KEYWORDS =
Keywords whose value is an array of subschemas.
%w[allOf anyOf oneOf prefixItems].freeze
- DATA_KEYWORDS =
Keywords whose value is data, not schema: never walked, never re-keyed.
%w[enum const default examples].freeze
- DIALECT_KEYWORDS =
Keywords that exist only in some dialects, with the dialects that define them. A keyword absent from this table exists in every supported dialect. Under a dialect that does not define a keyword the keyword is an unknown one: ignored, never shape-checked, walked, evaluated or reported as unsupported.
{ 'prefixItems' => [DEFAULT_DIALECT], '$dynamicRef' => [DEFAULT_DIALECT], '$dynamicAnchor' => [DEFAULT_DIALECT], '$recursiveRef' => [DRAFT_2019_09], '$recursiveAnchor' => [DRAFT_2019_09], '$vocabulary' => [DEFAULT_DIALECT, DRAFT_2019_09], 'additionalItems' => [DRAFT_2019_09, DRAFT_07], 'dependencies' => [DRAFT_07], 'dependentSchemas' => [DEFAULT_DIALECT, DRAFT_2019_09], 'dependentRequired' => [DEFAULT_DIALECT, DRAFT_2019_09], 'unevaluatedItems' => [DEFAULT_DIALECT, DRAFT_2019_09], 'unevaluatedProperties' => [DEFAULT_DIALECT, DRAFT_2019_09], 'contentSchema' => [DEFAULT_DIALECT, DRAFT_2019_09], 'deprecated' => [DEFAULT_DIALECT, DRAFT_2019_09], 'minContains' => [DEFAULT_DIALECT, DRAFT_2019_09], 'maxContains' => [DEFAULT_DIALECT, DRAFT_2019_09], '$anchor' => [DEFAULT_DIALECT, DRAFT_2019_09], '$defs' => [DEFAULT_DIALECT, DRAFT_2019_09] }.freeze
- ANCHOR_NAME_BY_DIALECT =
Plain-name fragment syntax, per dialect: 2020-12 Core Section 8.2.2 admits a leading underscore and no colon, while 2019-09 Core Section 8.2.3 and draft-07 Core Section 8.2.3 admit a colon and require a leading letter. A name legal in one dialect is not one in the other, and the dialect in force at the resource decides.
{ DEFAULT_DIALECT => /\A[A-Za-z_][-A-Za-z0-9._]*\z/, DRAFT_2019_09 => /\A[A-Za-z][-A-Za-z0-9.:_]*\z/, DRAFT_07 => /\A[A-Za-z][-A-Za-z0-9.:_]*\z/ }.freeze
- ANCHOR_NAME =
The 2020-12 plain-name syntax, the default when no dialect is known.
ANCHOR_NAME_BY_DIALECT.fetch(DEFAULT_DIALECT)
- MAX_STRUCTURAL_OBJECTS =
Structural elements (schemas, the keyword maps holding them, and array members — boolean subschemas included) a schema document may contain before it is rejected unread. A usable schema has at most MAX_SUBSCHEMAS subschemas, each with a handful of keyword maps at most, so this bound only ever stops documents the preflight would reject anyway.
MAX_SUBSCHEMAS * 4
- UNRESOLVED =
Marker for a pointer that does not resolve (nil is a valid schema value position, so it cannot serve as the marker).
Object.new.freeze
Constants included from UriReferences
Constants included from Shapes
Shapes::ANNOTATION_SHAPES, Shapes::ASSERTION_SHAPES, Shapes::JSON_TYPE_NAMES, Shapes::KEYWORD_SHAPES, Shapes::VALUE_SHAPES
Constants included from Evaluation
Evaluation::COMPOSITION_STEPS, Evaluation::REFERENCE_KEYWORDS
Constants included from EcmaPatterns
EcmaPatterns::ECMA_ANCHORS, EcmaPatterns::ECMA_ANY_CLASS, EcmaPatterns::ECMA_DOT, EcmaPatterns::ECMA_EMPTY_CLASS, EcmaPatterns::ECMA_GROUP_OPENERS, EcmaPatterns::ECMA_KEPT_ESCAPES, EcmaPatterns::ECMA_NON_BOUNDARY, EcmaPatterns::ECMA_SPACE_MEMBERS, EcmaPatterns::ECMA_WORD, EcmaPatterns::ECMA_WORD_BOUNDARY, EcmaPatterns::TRANSLATION_CHECK_INTERVAL
Class Method Summary collapse
-
.admit_schema?(schema, depth, counter, problems) ⇒ Boolean
Account for a schema (object or boolean) about to be walked: once per object, and within the subschema and nesting bounds — a boolean is a subschema too and obeys the depth bound.
-
.anchor_index_problems(root, resolver) ⇒ Array<String>
Problems the anchor index reveals: anchor names must be unique within a schema resource (JSON Schema 2020-12 Core Section 8.2.2) and a resource URI unique within the document (Section 9.1.2) — either declared twice would bind a reference to whichever declaration the walk met first — and an index that stopped at its bound before reaching every object leaves references and resource dialects undecided, so either makes the schema unusable.
-
.anchor_name?(name, dialect) ⇒ Boolean
Whether the dialect admits the name as a plain name.
-
.budget_exhausted?(deadline) ⇒ Boolean
Whether the validation-wide deadline has passed.
-
.charge_boolean_target(hop, root, from, counter, problems) ⇒ void
Count a boolean a reference reaches toward the subschema bound, once per distinct position (resource and decoded pointer): a document may not hide thousands of applied schemas behind pointer-addressable booleans.
-
.check_deadline(ctx) ⇒ Object
Consult the validation-wide deadline without spending the node-visit allowance.
-
.check_dynamic_refs(walk, schema, depth, dialect) ⇒ Array<Array>
The dynamic references of a schema object.
-
.check_keyword_shapes(walk, schema, dialect) ⇒ void
Every check a schema object's own keywords get.
-
.check_normalized(root, counter = {}) ⇒ Array<String>
SchemaValidator.check_schema on an already normalized schema.
-
.check_ref(walk, schema, depth, dialect, keyword = '$ref') ⇒ Array<Array>
Check the
$refof a schema object and queue what it reaches: a pointer may lead into a bag the dialect does not walk ($defsunder draft-07), and what a reference applies must be usable too. -
.check_schema(schema, counter = {}, deadline: nil) ⇒ Array<String>
Check that a schema can be used at all: it is an object or a boolean, its dialect is supported, it stays within the resource bounds, and every
$refresolves inside the document to a schema (a network or otherwise external reference is never dereferenced and makes the schema unusable rather than permissive). -
.clip(text) ⇒ String
Bound a piece of peer-derived text destined for a message.
-
.clip_value(value) ⇒ String
A short rendering of a value for a message that never inspects a large value whole.
-
.count_errors(ctx, errors, already_counted: 0) ⇒ Array<String>
Account for produced errors against MAX_ERRORS.
-
.count_visit(ctx) ⇒ Object
Account for one node visit (boolean schemas included, so a huge array under
items: truestill runs into the bounds). -
.each_definition(schema, dialect = nil, &block) ⇒ void
Yield the reusable schemas under the definition bag(s) the dialect defines:
definitionseverywhere (2020-12 Validation Appendix A keeps it as the deprecated spelling of$defs) and$defsin 2019-09 and 2020-12; both when no dialect is given. -
.each_subschema(schema, dialect = nil, &block) ⇒ void
Yield every subschema directly under a schema object (skipping the keywords the dialect does not define; nil applies no dialect).
-
.entered_scope?(ctx, schema) ⇒ Boolean
Enter the schema resource a subschema belongs to, when applying it leaves the resource in force.
-
.follow_ref_chain(ref, root, dialect, resolver, from) ⇒ String?
The problem, if any; each resolved target is yielded.
-
.integer?(data) ⇒ Boolean
Whether a value is a JSON Schema integer.
-
.json_type(data) ⇒ String
The JSON type name of a Ruby value (for error messages).
-
.keyword_known?(keyword, dialect) ⇒ Boolean
Whether the dialect defines the keyword.
- .leave_scope(ctx, entered) ⇒ void
-
.node_dialect(schema, ctx) ⇒ String?
The dialect in force at a schema object during validation: the one recorded for its resource by the anchor index (built once per validation), else the root's.
-
.normalize_schema(schema, deadline: nil) ⇒ Object
A bounded, string-keyed copy of a schema (booleans pass through).
-
.queue_definitions(walk, schema, depth, dialect, referenced) ⇒ void
Queue the reusable schemas beside a draft-07
$ref, then what the reference reaches (walked first, so pushed last). - .queue_positions(walk, positions, referenced) ⇒ void
-
.queue_subschemas(walk, schema, depth, dialect, referenced) ⇒ void
Queue every subschema position under a schema object, then what its reference reaches.
- .record_embedded_dialect_problem(walk, schema, problem) ⇒ void
-
.ref_chain_problem(ref, root, dialect, resolver, from:, &block) ⇒ String?
Follow a local reference (and the references it leads to) at preflight: every hop must resolve to a schema, the chain must not cycle, and it must stay within MAX_REF_DEPTH.
-
.referenced_target_position(walk, target, hop, from, depth, dialect) ⇒ Array?
The position a reference's target is walked at: under its own resource's dialect and at its own lexical depth.
-
.schema_value?(value) ⇒ Boolean
Whether the value is a schema (object or boolean).
-
.subschemas_under(keyword, value) ⇒ Array<Object>
The subschema positions one keyword holds.
-
.type_match?(type, data) ⇒ Boolean
Whether a value matches a JSON Schema type name.
-
.unsupported_dialect(schema) ⇒ String?
The dialect a schema declares (at its root or at an embedded resource root) that this validator does not implement.
-
.validate(data, schema, path: '#', deadline: nil) ⇒ Array<String>
Validate data against a schema.
-
.validate_child(data, schema, path, ctx) ⇒ Array<String>
Validate a child of the value being validated — an array item or a property value — against the subschema for it.
-
.validate_enum(data, schema, path) ⇒ Array<String>
Validate enum/const membership.
-
.validate_node(data, schema, path, ctx, ref_depth) ⇒ Array<String>
Validate one value against one (sub)schema, and against every schema that one leads to for the same value: a
$refhop, an allOf/anyOf/oneOf/not/if branch, a then/else. -
.validate_type(data, type, path) ⇒ Array<String>
Validate the JSON type of a value.
-
.walk_position(walk, schema, depth, dialect) ⇒ void
Walk one schema position and queue the positions it leads to.
-
.walk_schema(schema, root, depth, counter, problems, dialect = ) ⇒ void
Walk every subschema position, checking bounds and references.
Methods included from Dialects
canonical_dialect, dialect, embedded_dialect, embedded_dialect_problem, supported_dialect?
Methods included from Normalization
charge_structure, deep_stringify, list_member?, member_mode, stringify_member
Methods included from UriReferences
compose_uri, dot_segment_step, merge_paths, merge_relative_uri, merge_uri, parse_uri_reference, remove_dot_segments, strip_leading_dots
Methods included from References
adopt_reached_target, adopt_step, anchor_index, anchor_names, decode_component, decoded_fragment, dynamic_binding, each_foreign_definition, each_walked_position, enter_resource, external_ref?, index_positions, indexed_dialect, lexical_depths, normalized_copy, outermost_dynamic_anchor, outermost_recursive_anchor, pointer_child, pointer_origin, pointer_position, pointer_step, pointer_step_mode, pointer_tokens, record_anchor_names, referenced_position_depth, register_resource_base, resolve_adopted_pointer, resolve_reference, resource_root?, resource_start?, retarget_reference
Methods included from Shapes
absolute_uri?, all_schemas?, applicator_shape_problem, array_value?, assertion_shape_problem, boolean_value?, check_applicator_shapes, check_assertion_shapes, check_core_keyword_shapes, check_exclusive_bounds, check_identifier_shapes, check_pattern_shapes, dependencies_shape_problem, dependent_required_shape_problem, id_shape_problem, non_negative_integer?, number_value?, pattern_shape_problem, positive_number?, property_names?, shape_requirement, string_value?, type_shape_problem, uri_reference?, vocabulary_map?
Methods included from KeywordScan
queue_scan, referenced_position, scan_position, scan_positions, unsupported_keywords
Methods included from Composition
contains_max, contains_min, partial_keywords?, pattern_matches?, property_present?
Methods included from Evaluation
all_of_branch, applied_references, apply_keywords, apply_references, apply_unevaluated, branch_verdict, branch_verdicts, compose, compose_all_of, compose_any_of, compose_conditional, compose_dependent_schemas, compose_not, compose_one_of, dependency_branch, finish_node, merge_passed, reads_annotations?, ref_problem, ref_target, resume_conditional, start_node, triggered_dependencies, unconditional_conditional, unusable_ref?
Methods included from Scalars
multiple_of?, pattern_budget_remaining, validate_number, validate_pattern, validate_string
Methods included from EcmaPatterns
class_member, class_source, close_repeat_group, copy_brace, copy_character_class, copy_ecma_token, copy_escape, copy_group_opener, copy_named_reference, copy_numeric_escape, copy_plain_token, copy_quantifier, copy_unicode_escape, count_capture_groups, ecma_escape, ecma_regexp, ecma_source, emit, generated_name_prefix, group_name_at, group_name_for, hex_code_unit, note_translation_progress, open_repeat_group, quantifier_at?, reject_repeated_reference, repeated_bounds?, repeated_capture_groups
Methods included from InputRequirements
input_requirements, position_dialect, queue_referenced, read_requirement_positions, read_requirements
Methods included from Instances
comparable_value, contains_annotates?, count_contains_matches, exact_number, note_evaluated_items, property_errors, property_name_errors, speculatively, unevaluated_errors, unevaluated_item_errors, unevaluated_property_errors, validate_array, validate_contains, validate_dependent_required, validate_items, validate_named_properties, validate_object, validate_other_properties, validate_property_counts, validate_unique_items
Class Method Details
.admit_schema?(schema, depth, counter, problems) ⇒ Boolean
Account for a schema (object or boolean) about to be walked: once per object, and within the subschema and nesting bounds — a boolean is a subschema too and obeys the depth bound.
506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 |
# File 'lib/mcp_client/schema_validator.rb', line 506 def self.admit_schema?(schema, depth, counter, problems) return false if schema.is_a?(Hash) && counter[:walked].key?(schema) counter[:walked][schema] = true if schema.is_a?(Hash) counter[:count] += 1 if counter[:count] > MAX_SUBSCHEMAS problems << "schema has more than #{MAX_SUBSCHEMAS} subschemas" return false end if depth > MAX_SCHEMA_DEPTH problems << "schema nesting depth exceeds #{MAX_SCHEMA_DEPTH}" return false end schema.is_a?(Hash) end |
.anchor_index_problems(root, resolver) ⇒ Array<String>
Problems the anchor index reveals: anchor names must be unique within a schema resource (JSON Schema 2020-12 Core Section 8.2.2) and a resource URI unique within the document (Section 9.1.2) — either declared twice would bind a reference to whichever declaration the walk met first — and an index that stopped at its bound before reaching every object leaves references and resource dialects undecided, so either makes the schema unusable.
382 383 384 385 386 387 388 389 390 391 392 393 394 395 |
# File 'lib/mcp_client/schema_validator.rb', line 382 def self.anchor_index_problems(root, resolver) index = (resolver[:anchors] ||= anchor_index(root, resolver[:dialect])) problems = index[:duplicates].map do |name| "anchor #{clip(name.inspect)} is declared more than once in a schema resource" end problems.concat(index[:duplicate_ids].uniq.map do |base| "schema resource #{clip(base.inspect)} is declared more than once" end) if index[:truncated] problems << 'schema has too many or too deeply nested objects to index its anchors ' \ "(more than #{MAX_SUBSCHEMAS}, or deeper than #{MAX_SCHEMA_DEPTH})" end problems end |
.anchor_name?(name, dialect) ⇒ Boolean
Returns whether the dialect admits the name as a plain name.
262 263 264 |
# File 'lib/mcp_client/schema_validator.rb', line 262 def self.anchor_name?(name, dialect) name.is_a?(String) && name.match?(ANCHOR_NAME_BY_DIALECT.fetch(dialect, ANCHOR_NAME)) end |
.budget_exhausted?(deadline) ⇒ Boolean
Returns whether the validation-wide deadline has passed.
903 904 905 |
# File 'lib/mcp_client/schema_validator.rb', line 903 def self.budget_exhausted?(deadline) deadline && Process.clock_gettime(Process::CLOCK_MONOTONIC) >= deadline end |
.charge_boolean_target(hop, root, from, counter, problems) ⇒ void
This method returns an undefined value.
Count a boolean a reference reaches toward the subschema bound, once per distinct position (resource and decoded pointer): a document may not hide thousands of applied schemas behind pointer-addressable booleans.
622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 |
# File 'lib/mcp_client/schema_validator.rb', line 622 def self.charge_boolean_target(hop, root, from, counter, problems) index = (counter[:anchors] ||= anchor_index(root, counter[:dialect])) resource, hop = pointer_origin(index, root, hop, from) return unless resource tokens = pointer_tokens(hop) return unless tokens key = [resource.object_id, tokens] seen = (counter[:boolean_targets] ||= {}) return if seen.key?(key) seen[key] = true counter[:count] += 1 problems << "schema has more than #{MAX_SUBSCHEMAS} subschemas" if counter[:count] > MAX_SUBSCHEMAS end |
.check_deadline(ctx) ⇒ Object
Consult the validation-wide deadline without spending the node-visit allowance. A loop over the instance's own members (an object's properties, an array's items) is as long as the peer made it and may decide a member without visiting a node for it, so the loop itself has to reach a checkpoint: otherwise the budget is only consulted between the nodes such a sweep happens to visit, and a wide enough instance runs past it — reported, inside a speculative branch, as a pass.
885 886 887 |
# File 'lib/mcp_client/schema_validator.rb', line 885 def self.check_deadline(ctx) raise Aborted, 'validation time budget exhausted' if budget_exhausted?(ctx.deadline) end |
.check_dynamic_refs(walk, schema, depth, dialect) ⇒ Array<Array>
The dynamic references of a schema object. One that names no dynamic
anchor is the plain reference it resolves to and is checked (and what
it reaches queued) exactly as a $ref is; a dynamic one is not
evaluated, but one pointing outside the document would need a fetch,
which never happens: the schema is unusable.
645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 |
# File 'lib/mcp_client/schema_validator.rb', line 645 def self.check_dynamic_refs(walk, schema, depth, dialect) DYNAMIC_REFERENCE_KEYWORDS.flat_map do |keyword| next [] unless schema.key?(keyword) && keyword_known?(keyword, dialect) ref = schema[keyword] # A reference that is not a URI reference at all names nothing: it # is not silently ignored (a malformed keyword is not an absent one). unless ref.is_a?(String) walk.problems << "#{keyword} must be a string, got #{json_type(ref)}" next [] end if external_ref?(ref, walk.root, walk.counter[:dialect], walk.counter, from: schema) walk.problems << "external #{keyword} #{clip(ref.inspect)} is not dereferenced" next [] end check_ref(walk, schema, depth, dialect, keyword) end end |
.check_keyword_shapes(walk, schema, dialect) ⇒ void
This method returns an undefined value.
Every check a schema object's own keywords get.
489 490 491 492 493 494 495 496 497 498 499 500 |
# File 'lib/mcp_client/schema_validator.rb', line 489 def self.check_keyword_shapes(walk, schema, dialect) problems = walk.problems if dialect == DEFAULT_DIALECT && schema['items'].is_a?(Array) problems << 'items must be a schema in JSON Schema 2020-12 (positional schemas go in prefixItems)' end check_applicator_shapes(schema, dialect, problems) check_assertion_shapes(schema, dialect, problems) check_exclusive_bounds(schema, dialect, problems) check_identifier_shapes(schema, dialect, problems) check_core_keyword_shapes(schema, dialect, problems) check_pattern_shapes(schema, dialect, problems, walk.counter[:deadline]) end |
.check_normalized(root, counter = {}) ⇒ Array<String>
check_schema on an already normalized schema.
350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 |
# File 'lib/mcp_client/schema_validator.rb', line 350 def self.check_normalized(root, counter = {}) return [] if [true, false].include?(root) return ["schema must be an object or a boolean, got #{json_type(root)}"] unless root.is_a?(Hash) declared = dialect(root) return ['$schema must be a non-empty string naming the dialect'] if declared.nil? unless supported_dialect?(declared) counter[:unsupported_dialect] = declared return ["schema dialect #{clip(declared.inspect)} is not supported " \ "(supported: #{SUPPORTED_DIALECTS.join(', ')})"] end problems = [] counter.update(count: 0, dialect: canonical_dialect(declared), walked: {}.compare_by_identity, depths: lexical_depths(root, canonical_dialect(declared))) walk_schema(root, root, 0, counter, problems) exhausted = problems.empty? && budget_exhausted?(counter[:deadline]) problems << 'validation aborted: validation time budget exhausted during the schema check' if exhausted problems.concat(anchor_index_problems(root, counter)) if problems.empty? problems.uniq end |
.check_ref(walk, schema, depth, dialect, keyword = '$ref') ⇒ Array<Array>
Check the $ref of a schema object and queue what it reaches: a
pointer may lead into a bag the dialect does not walk ($defs under
draft-07), and what a reference applies must be usable too.
567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 |
# File 'lib/mcp_client/schema_validator.rb', line 567 def self.check_ref(walk, schema, depth, dialect, keyword = '$ref') root = walk.root counter = walk.counter problems = walk.problems ref = schema[keyword] unless ref.is_a?(String) problems << "#{keyword} must be a string, got #{json_type(ref)}" return [] end if external_ref?(ref, root, counter[:dialect], counter, from: schema) problems << "external #{keyword} #{clip(ref.inspect)} is not dereferenced (only references inside the " \ 'schema document are resolved; network $ref resolution is disabled)' return [] end positions = [] problem = ref_chain_problem(ref, root, counter[:dialect], counter, from: schema) do |target, hop, from| position = referenced_target_position(walk, target, hop, from, depth, dialect) positions << position if position end problems << problem.sub('$ref', keyword) if problem positions end |
.check_schema(schema, counter = {}, deadline: nil) ⇒ Array<String>
Check that a schema can be used at all: it is an object or a boolean,
its dialect is supported, it stays within the resource bounds, and
every $ref resolves inside the document to a schema (a network or
otherwise external reference is never dereferenced and makes the
schema unusable rather than permissive).
The check runs under a deadline like a validation does: the schema is
the peer's, and reading it (translating and compiling its patterns
above all) costs what the peer's text costs.
311 312 313 314 315 316 317 318 |
# File 'lib/mcp_client/schema_validator.rb', line 311 def self.check_schema(schema, counter = {}, deadline: nil) counter[:deadline] = deadline || (Process.clock_gettime(Process::CLOCK_MONOTONIC) + PATTERN_MATCH_TIMEOUT) check_normalized(normalize_schema(schema, deadline: counter[:deadline]), counter) rescue TooLarge => e [e.] rescue Aborted => e ["validation aborted: #{e.}"] end |
.clip(text) ⇒ String
Bound a piece of peer-derived text destined for a message.
985 986 987 988 |
# File 'lib/mcp_client/schema_validator.rb', line 985 def self.clip(text) text = text.to_s text.length > MAX_VALUE_INSPECT ? "#{text[0, MAX_VALUE_INSPECT]}..." : text end |
.clip_value(value) ⇒ String
A short rendering of a value for a message that never inspects a large value whole.
994 995 996 997 998 999 1000 1001 |
# File 'lib/mcp_client/schema_validator.rb', line 994 def self.clip_value(value) case value when String then clip(value[0, MAX_VALUE_INSPECT].inspect) when Array then "array(#{value.length} items)" when Hash then "object(#{value.length} keys)" else clip(value.inspect) end end |
.count_errors(ctx, errors, already_counted: 0) ⇒ Array<String>
Account for produced errors against MAX_ERRORS. Errors inside a speculative branch are a verdict, not output, and do not count.
893 894 895 896 897 898 899 900 |
# File 'lib/mcp_client/schema_validator.rb', line 893 def self.count_errors(ctx, errors, already_counted: 0) return errors if ctx.speculative.positive? ctx.errors += errors.size - already_counted raise Aborted, "more than #{MAX_ERRORS} validation errors (output truncated)" if ctx.errors > MAX_ERRORS errors end |
.count_visit(ctx) ⇒ Object
Account for one node visit (boolean schemas included, so a huge array
under items: true still runs into the bounds).
870 871 872 873 874 875 |
# File 'lib/mcp_client/schema_validator.rb', line 870 def self.count_visit(ctx) ctx.visits += 1 raise Aborted, "more than #{MAX_NODE_VISITS} schema nodes visited" if ctx.visits > MAX_NODE_VISITS check_deadline(ctx) end |
.each_definition(schema, dialect = nil, &block) ⇒ void
This method returns an undefined value.
Yield the reusable schemas under the definition bag(s) the dialect
defines: definitions everywhere (2020-12 Validation Appendix A keeps
it as the deprecated spelling of $defs) and $defs in 2019-09 and
2020-12; both when no dialect is given.
555 556 557 558 559 560 561 |
# File 'lib/mcp_client/schema_validator.rb', line 555 def self.each_definition(schema, dialect = nil, &block) %w[$defs definitions].each do |keyword| next unless keyword_known?(keyword, dialect) schema[keyword].each_value(&block) if schema[keyword].is_a?(Hash) end end |
.each_subschema(schema, dialect = nil, &block) ⇒ void
This method returns an undefined value.
Yield every subschema directly under a schema object (skipping the keywords the dialect does not define; nil applies no dialect).
525 526 527 528 529 530 531 |
# File 'lib/mcp_client/schema_validator.rb', line 525 def self.each_subschema(schema, dialect = nil, &block) schema.each do |keyword, value| next if DATA_KEYWORDS.include?(keyword) || !keyword_known?(keyword, dialect) subschemas_under(keyword, value).each(&block) end end |
.entered_scope?(ctx, schema) ⇒ Boolean
Enter the schema resource a subschema belongs to, when applying it leaves the resource in force. A schema of the resource already innermost adds nothing a dynamic reference could read, so it is not pushed and its application has nothing to pop.
821 822 823 824 825 826 827 828 829 830 |
# File 'lib/mcp_client/schema_validator.rb', line 821 def self.entered_scope?(ctx, schema) return false unless schema.is_a?(Hash) && ctx.scope ctx.anchors ||= anchor_index(ctx.root, ctx.dialect) resource = ctx.anchors[:resources][schema] || ctx.root return false if ctx.scope.last.equal?(resource) ctx.scope.push(resource) true end |
.follow_ref_chain(ref, root, dialect, resolver, from) ⇒ String?
Returns the problem, if any; each resolved target is yielded.
680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 |
# File 'lib/mcp_client/schema_validator.rb', line 680 def self.follow_ref_chain(ref, root, dialect, resolver, from) # A cycle is a chain returning to a schema it already reached: hops # are told apart by where they land, not by their fragment text, # since the same fragment means something else in another resource. seen = {}.compare_by_identity current = ref loop do return "$ref chain #{clip(ref.inspect)} exceeds #{MAX_REF_DEPTH} hops" if seen.size >= MAX_REF_DEPTH target = resolve_reference(root, current, dialect, resolver, from: from) return "unresolvable local $ref #{clip(current.inspect)}" if target.equal?(UNRESOLVED) return "$ref #{clip(current.inspect)} does not point at a schema" unless schema_value?(target) return "$ref chain #{clip(ref.inspect)} cycles" if target.is_a?(Hash) && seen.key?(target) seen[target] = true if target.is_a?(Hash) yield target, current, from return nil unless target.is_a?(Hash) && target.key?('$ref') from = target current = target['$ref'] return "$ref must be a string, got #{json_type(current)}" unless current.is_a?(String) if external_ref?(current, root, dialect, resolver, from: from) return "external $ref #{clip(current.inspect)} is not dereferenced" end end end |
.integer?(data) ⇒ Boolean
Whether a value is a JSON Schema integer. Per JSON Schema 2020-12 a number with a zero fractional part (e.g. 2.0) is a valid integer.
941 942 943 944 945 946 |
# File 'lib/mcp_client/schema_validator.rb', line 941 def self.integer?(data) return true if data.is_a?(Integer) return false unless data.is_a?(Numeric) (data % 1).zero? end |
.json_type(data) ⇒ String
The JSON type name of a Ruby value (for error messages).
951 952 953 954 955 956 957 958 959 960 961 962 |
# File 'lib/mcp_client/schema_validator.rb', line 951 def self.json_type(data) case data when nil then 'null' when true, false then 'boolean' when Integer then 'integer' when Numeric then 'number' when String then 'string' when Array then 'array' when Hash then 'object' else data.class.name end end |
.keyword_known?(keyword, dialect) ⇒ Boolean
Returns whether the dialect defines the keyword.
269 270 271 |
# File 'lib/mcp_client/schema_validator.rb', line 269 def self.keyword_known?(keyword, dialect) dialect.nil? || !DIALECT_KEYWORDS.key?(keyword) || DIALECT_KEYWORDS[keyword].include?(dialect) end |
.leave_scope(ctx, entered) ⇒ void
This method returns an undefined value.
835 836 837 |
# File 'lib/mcp_client/schema_validator.rb', line 835 def self.leave_scope(ctx, entered) ctx.scope.pop if entered && ctx.scope end |
.node_dialect(schema, ctx) ⇒ String?
The dialect in force at a schema object during validation: the one recorded for its resource by the anchor index (built once per validation), else the root's.
862 863 864 865 |
# File 'lib/mcp_client/schema_validator.rb', line 862 def self.node_dialect(schema, ctx) ctx.anchors ||= anchor_index(ctx.root, ctx.dialect) indexed_dialect(schema, ctx) || ctx.dialect end |
.normalize_schema(schema, deadline: nil) ⇒ Object
A bounded, string-keyed copy of a schema (booleans pass through).
341 342 343 |
# File 'lib/mcp_client/schema_validator.rb', line 341 def self.normalize_schema(schema, deadline: nil) schema.is_a?(Hash) ? deep_stringify(schema, 0, { objects: 0, deadline: deadline }) : schema end |
.queue_definitions(walk, schema, depth, dialect, referenced) ⇒ void
This method returns an undefined value.
Queue the reusable schemas beside a draft-07 $ref, then what the
reference reaches (walked first, so pushed last).
466 467 468 469 470 |
# File 'lib/mcp_client/schema_validator.rb', line 466 def self.queue_definitions(walk, schema, depth, dialect, referenced) positions = [] each_definition(schema, dialect) { |sub| positions << [sub, depth + 1, dialect] } queue_positions(walk, positions, referenced) end |
.queue_positions(walk, positions, referenced) ⇒ void
This method returns an undefined value.
482 483 484 485 |
# File 'lib/mcp_client/schema_validator.rb', line 482 def self.queue_positions(walk, positions, referenced) walk.pending.concat(positions.reverse) walk.pending.concat(referenced.reverse) end |
.queue_subschemas(walk, schema, depth, dialect, referenced) ⇒ void
This method returns an undefined value.
Queue every subschema position under a schema object, then what its reference reaches.
475 476 477 478 479 |
# File 'lib/mcp_client/schema_validator.rb', line 475 def self.queue_subschemas(walk, schema, depth, dialect, referenced) positions = [] each_subschema(schema, dialect) { |sub| positions << [sub, depth + 1, dialect] } queue_positions(walk, positions, referenced) end |
.record_embedded_dialect_problem(walk, schema, problem) ⇒ void
This method returns an undefined value.
457 458 459 460 461 |
# File 'lib/mcp_client/schema_validator.rb', line 457 def self.(walk, schema, problem) declared = dialect(schema) walk.counter[:unsupported_dialect] ||= declared if declared && !supported_dialect?(declared) walk.problems << problem end |
.ref_chain_problem(ref, root, dialect, resolver, from:, &block) ⇒ String?
Follow a local reference (and the references it leads to) at preflight: every hop must resolve to a schema, the chain must not cycle, and it must stay within MAX_REF_DEPTH. Once the whole chain checks out, every target it reached is yielded so the caller can preflight what the reference applies.
672 673 674 675 676 677 |
# File 'lib/mcp_client/schema_validator.rb', line 672 def self.ref_chain_problem(ref, root, dialect, resolver, from:, &block) targets = [] problem = follow_ref_chain(ref, root, dialect, resolver, from) { |*hop| targets << hop } targets.each { |target, used, origin| block.call(target, used, origin) } if problem.nil? && block problem end |
.referenced_target_position(walk, target, hop, from, depth, dialect) ⇒ Array?
The position a reference's target is walked at: under its own resource's dialect and at its own lexical depth. A boolean target needs no preflight and is charged once per distinct position (not per reference), and that position still obeys the depth bound.
596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 |
# File 'lib/mcp_client/schema_validator.rb', line 596 def self.referenced_target_position(walk, target, hop, from, depth, dialect) counter = walk.counter position, opaque, visited = pointer_position(hop, walk.root, counter[:dialect], counter, from) unless target.is_a?(Hash) walk.problems << "schema nesting depth exceeds #{MAX_SCHEMA_DEPTH}" if position && position > MAX_SCHEMA_DEPTH # A boolean in a position the walk visits was admitted there; one # the walk never reaches (behind an opaque keyword, or beside a # draft-07 $ref) is charged here, once per position. charge_boolean_target(hop, walk.root, from, counter, walk.problems) unless visited return nil end lexical = (counter[:depths] || {})[target] target_depth = if opaque [position, lexical].compact.max || (depth + 1) else lexical || position || (depth + 1) end [target, target_depth, indexed_dialect(target, counter) || dialect] end |
.schema_value?(value) ⇒ Boolean
Returns whether the value is a schema (object or boolean).
709 710 711 |
# File 'lib/mcp_client/schema_validator.rb', line 709 def self.schema_value?(value) value.is_a?(Hash) || value == true || value == false end |
.subschemas_under(keyword, value) ⇒ Array<Object>
The subschema positions one keyword holds.
535 536 537 538 539 540 541 542 543 544 545 546 547 548 |
# File 'lib/mcp_client/schema_validator.rb', line 535 def self.subschemas_under(keyword, value) if SUBSCHEMA_KEYWORDS.include?(keyword) value.is_a?(Array) ? value : [value] elsif keyword == 'dependencies' # Property-name arrays are data; only the schema entries are walked. value.is_a?(Hash) ? value.values.select { |v| schema_value?(v) } : [] elsif SUBSCHEMA_MAP_KEYWORDS.include?(keyword) value.is_a?(Hash) ? value.values : [] elsif SUBSCHEMA_ARRAY_KEYWORDS.include?(keyword) value.is_a?(Array) ? value : [] else [] end end |
.type_match?(type, data) ⇒ Boolean
Whether a value matches a JSON Schema type name. Unknown type names are not enforced (returns true).
924 925 926 927 928 929 930 931 932 933 934 935 |
# File 'lib/mcp_client/schema_validator.rb', line 924 def self.type_match?(type, data) case type when 'object' then data.is_a?(Hash) when 'array' then data.is_a?(Array) when 'string' then data.is_a?(String) when 'boolean' then data.equal?(true) || data.equal?(false) when 'null' then data.nil? when 'number' then data.is_a?(Numeric) when 'integer' then integer?(data) else true end end |
.unsupported_dialect(schema) ⇒ String?
The dialect a schema declares (at its root or at an embedded resource root) that this validator does not implement. MCP 2026-07-28 basic "Implementation Requirements": a client "MUST handle unsupported dialects gracefully by returning an appropriate error indicating the dialect is not supported", which a caller can only do once it can tell an unsupported dialect from every other reason a schema is unusable.
329 330 331 332 333 |
# File 'lib/mcp_client/schema_validator.rb', line 329 def self.unsupported_dialect(schema) state = {} check_schema(schema, state) state[:unsupported_dialect] end |
.validate(data, schema, path: '#', deadline: nil) ⇒ Array<String>
Validate data against a schema. An unusable schema (see check_schema) is reported as validation errors, never as a pass, and so is a validation that hit a resource bound. Schema and data hashes may use string or symbol keys.
726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 |
# File 'lib/mcp_client/schema_validator.rb', line 726 def self.validate(data, schema, path: '#', deadline: nil) # One deadline covers the entire validation, the copy of the peer's # document included. deadline ||= Process.clock_gettime(Process::CLOCK_MONOTONIC) + PATTERN_MATCH_TIMEOUT root = normalize_schema(schema, deadline: deadline) # `preflight` comes back holding the state the check read the schema # under; the check runs under the validation's deadline. problems = check_normalized(root, preflight = { deadline: deadline }) return problems.map { |problem| "#{path}: #{problem}" } unless problems.empty? # The index the preflight built is the one this validation reads (the # resources, dialects and anchors a schema was accepted under, and the # subtrees its references adopted): a second one would make what a # reference resolves to depend on the order this instance applies them. ctx = Context.new(root: root, deadline: deadline, dialect: canonical_dialect(dialect(root)), visits: 0, errors: 0, speculative: 0, anchors: preflight[:anchors], undecided: 0, depth: 0, scope: []) validate_node(data, root, path, ctx, 0) rescue TooLarge => e ["#{path}: #{e.}"] rescue Aborted => e ["#{path}: validation aborted: #{e.}"] rescue SystemStackError # A backstop, not a working bound: the `$ref` chain and the composition # branches a schema spends on one value are applied iteratively, so # what one instance level costs is fixed and MAX_NODE_DEPTH bounds the # stack. Should some stack still be smaller than that — the walk holds # no state outside `ctx` — the unwound stack leaves nothing behind, and # the validation ends the way an exhausted budget does, as one error, # never as a crash out of a tool call. ["#{path}: validation aborted: schema too deeply recursive for this stack"] end |
.validate_child(data, schema, path, ctx) ⇒ Array<String>
Validate a child of the value being validated — an array item or a
property value — against the subschema for it. This is the one step
that nests the instance, so it is where the walk's depth is counted and
where the $ref hop budget starts over (see validate_object).
849 850 851 852 853 854 855 856 |
# File 'lib/mcp_client/schema_validator.rb', line 849 def self.validate_child(data, schema, path, ctx) ctx.depth += 1 raise Aborted, "instance nested deeper than #{MAX_NODE_DEPTH}" if ctx.depth > MAX_NODE_DEPTH validate_node(data, schema, path, ctx, 0) ensure ctx.depth -= 1 end |
.validate_enum(data, schema, path) ⇒ Array<String>
Validate enum/const membership. Symbol- and string-keyed objects are compared as given on both sides (the schema's data values are never re-keyed).
971 972 973 974 975 976 977 978 979 980 |
# File 'lib/mcp_client/schema_validator.rb', line 971 def self.validate_enum(data, schema, path) errors = [] if schema['enum'].is_a?(Array) && !schema['enum'].include?(data) errors << "#{path}: value #{clip_value(data)} is not in enum #{clip_value(schema['enum'])}" end if schema.key?('const') && schema['const'] != data errors << "#{path}: value #{clip_value(data)} does not equal const #{clip_value(schema['const'])}" end errors end |
.validate_node(data, schema, path, ctx, ref_depth) ⇒ Array<String>
Validate one value against one (sub)schema, and against every schema
that one leads to for the same value: a $ref hop, an
allOf/anyOf/oneOf/not/if branch, a then/else. Applying another schema
to the same value does not nest the instance, so it is the $ref hop
budget and the visit cap that bound it; validate_child accounts for
the steps that do nest.
Those same-instance applications run on the pending list here rather
than on the Ruby stack (see Evaluation): a schema composed through N
$defs mixins applies N of them to every instance level, and spending
a frame on each would make the stack grow with the product of the
instance depth and the mixin count — so data a peer can legitimately
send would overflow the small stack of a transport's reader thread. A
frame is spent only on the step into a child value, which is what
MAX_NODE_DEPTH bounds.
781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 |
# File 'lib/mcp_client/schema_validator.rb', line 781 def self.validate_node(data, schema, path, ctx, ref_depth) pending = [] # The dynamic scope grows and shrinks with these applications, exactly # as it does with the Ruby frames a recursive validator would spend: # entering a schema of another resource enters that resource, and the # application that entered it is the one that leaves it. entered = [entered_scope?(ctx, schema)] step = start_node(data, schema, path, ctx, ref_depth) loop do if step[0] == :apply pending << step # A speculative application's errors are a verdict, not output, and # do not count toward MAX_ERRORS while it runs. ctx.speculative += 1 if step[3] entered << entered_scope?(ctx, step[1]) step = start_node(data, step[1], path, ctx, step[2], collecting: step[5]) next end errors = step[1] leave_scope(ctx, entered.pop) return errors if pending.empty? resumed = pending.pop ctx.speculative -= 1 if resumed[3] # The finished application hands back its verdict and what it # evaluated of the value, for the applicator that applied it. step = resumed[4].call(errors, step[2]) end ensure entered&.each { |was_entered| leave_scope(ctx, was_entered) } end |
.validate_type(data, type, path) ⇒ Array<String>
Validate the JSON type of a value.
912 913 914 915 916 917 |
# File 'lib/mcp_client/schema_validator.rb', line 912 def self.validate_type(data, type, path) types = (type.is_a?(Array) ? type : [type]).map(&:to_s) return [] if types.any? { |t| type_match?(t, data) } ["#{path}: expected type #{clip(types.join(' or '))}, got #{json_type(data)}"] end |
.walk_position(walk, schema, depth, dialect) ⇒ void
This method returns an undefined value.
Walk one schema position and queue the positions it leads to. The
queue is a stack and children are pushed in reverse, so a document is
read depth-first and in document order, exactly as a recursive walk
would read it — but a $ref chain hundreds of schemas long costs
queue entries rather than interpreter frames, so a shallow document
whose references chain deeply can no longer overflow the (small) stack
of the thread a transport reads on. A reference is followed before the
position's own subschemas, so it is pushed last.
422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 |
# File 'lib/mcp_client/schema_validator.rb', line 422 def self.walk_position(walk, schema, depth, dialect) return unless schema_value?(schema) # The positions are as many as the peer's document holds, and each # may cost a pattern's translation: the deadline is consulted before # every one, and a check past it stops there ({.check_normalized} # reports it). return if budget_exhausted?(walk.counter[:deadline]) return unless admit_schema?(schema, depth, walk.counter, walk.problems) if resource_root?(schema, dialect) && schema.key?('$schema') problem = (schema) return (walk, schema, problem) if problem dialect = (schema, dialect) elsif schema.key?('$schema') && !schema.equal?(walk.root) && !resource_start?(schema) # A `$schema` is read at a schema resource root only (JSON Schema # 2020-12 Core Section 8.1.1); anywhere else it is a malformed # keyword, not an absent one. walk.problems << '$schema is only allowed at a schema resource root (the document root, or a ' \ 'subschema declaring an $id)' return end referenced = schema.key?('$ref') ? check_ref(walk, schema, depth, dialect) : [] # draft-07: the $ref replaces its siblings, so the applicators next to # it are never applied and are not preflighted either; definitions are # a bag of reusable schemas, not applicators, and stay reachable # through references. return queue_definitions(walk, schema, depth, dialect, referenced) if dialect == DRAFT_07 && schema.key?('$ref') referenced.concat(check_dynamic_refs(walk, schema, depth, dialect)) check_keyword_shapes(walk, schema, dialect) queue_subschemas(walk, schema, depth, dialect, referenced) end |
.walk_schema(schema, root, depth, counter, problems, dialect = ) ⇒ void
This method returns an undefined value.
Walk every subschema position, checking bounds and references. A
schema object is walked once, however many positions or references
lead to it, so a recursive schema stays within the bounds. The dialect
follows the resource: an embedded resource declaring $schema is
walked under its own dialect.
403 404 405 406 407 408 409 410 411 |
# File 'lib/mcp_client/schema_validator.rb', line 403 def self.walk_schema(schema, root, depth, counter, problems, dialect = counter[:dialect]) walk = Walk.new(root: root, counter: counter, problems: problems, pending: [[schema, depth, dialect]]) until walk.pending.empty? return unless problems.empty? walk_position(walk, *walk.pending.pop) end end |