DIANA is an instance of IDL, IDL is a language/method for defining a data-structure [think “abstract data type” which is “implementable many different ways”], which produces (in the target language) the appropriate realization.
For example; in IDL there is no enumeration, however, an enumeration can be inferred via ‘empty’ classes, as per the tutorial (pg 8):
IDL does not have a special mechanism for creating enumeration types; instead, the strict class mechanism serves that purpose. If a class consists entirely of node types without attributes, it behaves much like an enumeration type. For example,
operation ::= plus | minus | times | divide;
plus =>; minus =>; times =>; divide =>;
defines class operation, which serves as the type of the “operation code” field of binary expressions.
and page 16, detailing how “Users can also gain more explicit control by using implementation notes.” — example:
A note with a name reference usually applies to a named type and all
attributes of the type; thus
operation ::= plus | minus | times | divide;
for operation use Enumeration;
says that type operation (a strict class type) should have an enumeration
type as its representation.
Thus, you could have the Ada-writer default to using Ada.Containers.Indefinite_Multiway_Tree instantiated against a base interface-type, or discriminated-record, or however you wish to represent them… then simply use the Tree’s stream attribute.
But who said anything about having it be a relational DB? It could (IMO should) be a graph DB, albeit with a ‘View’ for Relational, Hierarchical, Document, Object and/or whatever else DB-paradigm you need — is it going to be the super-fastest optimized-for-X’s-special-case! No, but we don’t necessarily need that.
A note on quick-and-dirty implementation of IDL: use the two assignments ::= and => to collect classes with Maps of Identifier → Vector of Identifier for each; if the class Identifier shows up in ::= and all of its “decedents” yield Empty_Vector in =>, you can use enumeration.
Function Enumerable( ID : Identifier; Classes, Attributes: Identifier_to_List ) return Boolean is
(for all X of Classes(ID) => Attribute(X) = Empty_Vector);
Warning: Very quick and dirty; there are cases this is wholly inappropriate for (akin to the “just split on comma!” ‘parsing’ of CSV), but perhaps suitable for tinkering or bootstrapping.
If bootstrapping [IDL in IDL, as noted in my previous post], then you could go with a “structure first” approach [that is, the graph that is the structure(s) of the IDL], using the quick-and-dirty to populate the graph, then build out a proper parser/serialization pair. (See Pair grammars, graph languages and string-to-graph translations [Alt-Link] — I disagree somewhat with the abstract: it is not String ↔ Graph that is of particular interest, but Token ↔ Graph; drop reading and tokenizing as concerns, considering them ‘already solved’.)
I would proffer that Byron’s Readington and Lexington, are decently good foundations for considering reading, and a good portion of tokenizing, solved — though I would say that the internals of Lexington.Token should be revised: using the 32 bits packed to the the record:
For Token use record
ID at 0 range 00..31;
Length at 0 range 32..63;
End record;
Where ID would, in the overlay-refinement be
For ID use record
Language at 00..07; -- 8-bits for the Language enumeration.
Lex_ID at 08..32; -- 24-bits for the Token-ID enumeration for the language.
End record;
reserving the “all-ones” ('Last) bit-value of each Language’s Token-ID for the unprocessed text from the reader, reserving the “all-zeroes” as a Null_Token. (Byron’s model has non-emittable tokens and uses them as temporary/intermediates in tokenizing; this makes things much easier afterwards: instead of dealing with ch_Colon [a raw/unprocessed colon] at the output, I have all the colon-containing items like semantic-separator ss_Assign [denoting the :=] in their own tokens, then we can grab all the colons-as-separators as ns_Colon, and thus ch_Colon appearing in the output of tokenization is in error.)
Once you regard reading/tokenizing as solved/separate parts you can then use what I’ve been calling the R1/R2 test:
- Obtain (Read in, Tokenize, Parse; or otherwise generate) some structure;
- Export the structure, this is “R1”, keep a copy;
- Read/Tokenize/Parse the export, this is R2;
- Compare R1 and R2, pass on identicality.
This testing works quite well for a “structure-first” approach and is essentially equivalent to testing seralize/deserialize round-tripping.