Binary encoding - 1. Avro
This is the first post in a series comparing different binary encoding technologies for messages stored in an event bus. After some research the technologies usable in a polyglot production environment are:
- Avro
- JSON with Message Pack and JSON schema.
- Protobuf
The first three posts present the following for each technology:
- The key differentiation factors
- The available tooling by showing the implementation of a hypothetical blog post comment creation event:
{
"id": "de2df598-9948-4988-b00a-a41c0e287398",
"message": "I’m sorry, Dave. I’m afraid I can’t do that.",
"locale": "en_US",
"author": {
"id": "78fc52a3-9f94-43c6-8eb5-9591e80b87e1",
"nickname": "HAL 9000",
},
"blog_id": "b4e05776-fca3-485e-be48-b1758cedd792",
"blog_title": "Binary encoding"
}
The event content is what we can expect in an Event-Carried State Transfer pattern
- The ecosystem: quality of documentation and tooling, available resources etc.
The final post compares the three technologies and tries to identify when each one should be used.
Let’s begin with Avro as it’s the least popular and thus should require more focus. The presentation below is quite detailed because of the documentation not being straightforward for this use case.
Key differentiation factors
Encoding algorithm
The key differentiating feature of Avro is its minimalist encoding algorithm:
- Types are not encoded
- Field names or indexes are not encoded
- Fields’ values are encoded in the order of the schema
Consequently: Consumers cannot decode, or even partially read, a message without the exact same version of the schema used by the producer to encode it.
Note: The Avro specification defines a resolution algorithm for when the consumer’s service is programmed to decode with an older version of the schema. To perform the resolution, both versions of the schema are required!
The Avro specification defines two different protocols for consumers to recover the producer’s schema:
- For large files with millions of records (eg: Hadoop), the schema is encoded within the payload.
- For database records stored one-by-one (eg: event bus), the schema is too much overhead. Avro’s specification defines a single object encoding algorithm composed of the fingerprint of the object’s schema.
Thus for this use case, a schema registry is required as consumers need to retrieve it from the parsed fingerprint.
Basically, the single object encoding algorithm is (copied from the documentation):
- A two-byte marker,
C3 01, to show that the message is Avro and uses this single-record format (version 1). - The 8-byte little-endian
CRC-64-AVROfingerprint of the object’s schema. - The Avro object encoded using Avro’s binary encoding.
Note: for the decoding process, Avro uses attribute
namein schema to match against code base structure (or classes) attribute name.
Schema
An Avro schema is a valid JSON object that defines one and only one custom type, a Record, Enum etc. A Record is the equivalent of a JSON object or a message in protobuf. Below is a bunch of noteworthy features:
- Fields are required by default. Optional fields can be declared using the enum or union type (eg:
{null, long}). - For schema evolution purposes:
- A default value can be optionally provided, only used when reading instances that lack the field.
- Aliases can be optionally provided as alternate names.
- The union type allows fields to have multiple types.
- To avoid name collisions, namespaces can be defined.
- Potentially useful logical types (date, UUID, timestamp, duration).
- Record definition can be nested in a parent record.
The next paragraph describes the available tooling through an example.
Available tooling
Schema management
Avro provides an IDL (Interface Description Language) to ease schema management. It’s less verbose than JSON and, above all, allows using types defined in other schemas. To generate a standalone JSON schema from IDL, the IDL tool duplicates the definitions of external types and nests them. These features are fundamental when maintaining a complex schema by factoring out type definitions.
The comment creation event below written in Avro IDL illustrates the process.
locale.avdl:
enum Locale {
en_US,
fr_FR,
zh_CN
}
comment.avdl:
import idl "locale.avdl";
record Author {
uuid id;
string nickname;
}
record Comment {
uuid id;
string message;
Locale locale;
Author author;
uuid blog_id;
string blog_title;
}
With the command idl2schemata one standalone schema .avsc file is created for each custom type. Below the content of Comment.avsc:
{
"type" : "record",
"name" : "Comment",
"fields" : [ {
"name" : "id",
"type" : {
"type" : "string",
"logicalType" : "uuid"
}
}, {
"name" : "message",
"type" : "string"
}, {
"name" : "locale",
"type" : {
"type" : "enum",
"name" : "Locale",
"symbols" : [ "en_US", "fr_FR", "zh_CN" ]
}
}, {
"name" : "author",
"type" : {
"type" : "record",
"name" : "Author",
"fields" : [ {
"name" : "id",
"type" : {
"type" : "string",
"logicalType" : "uuid"
}
}, {
"name" : "nickname",
"type" : "string"
} ]
}
}, {
"name" : "blog_id",
"type" : {
"type" : "string",
"logicalType" : "uuid"
}
}, {
"name" : "blog_title",
"type" : "string"
} ]
}
Note: As you can see, to have a standalone schema, records are copied and nested.
Avro IDL provides more features (eg: annotations, import JSON schema, protocol for RPC). But for this use case, it’s pretty much all we’ve got for schema management. Nothing is provided for the schema registry; an issue about implementing one is marked as “Won’t do” by the maintainers.
A minimal schema registry should have the following features:
- Expose a
HashMap<fingerprint, schema>for consumers. - Push to the HashMap for producers.
Then each team could privately maintain their schema in IDL format. To release a new version, they push standalone schema generated by the IDL tool as described above.
Note: The IDL tool is only available as a .jar
Let’s have a look at how to use the Namespace Avro schema in Rust using the official crate apache-avro.
Rust implementation
Good news! The encoding/decoding implementation is compatible with serde! Structure can be described as we’re used to with an implementation of the AvroSchema trait:
#[derive(Debug, Serialize, Deserialize)]
struct Comment {
id: Uuid,
message: String,
locale: Locale,
author: Author,
blog_id: Uuid,
blog_title: String,
}
impl AvroSchema for Comment {
fn get_schema() -> Schema {
Schema::parse_str(include_str!("../schemas/standalone/Comment.avsc"))
.expect("Invalid Comment Avro schema")
}
}
Note: Macros could be used to assert schema validity at compile time.
Then encoding a message is straightforward:
let payload = Comment {
id: "de2df598-9948-4988-b00a-a41c0e287398".parse().unwrap(),
message: "I’m sorry, Dave. I’m afraid I can’t do that.".to_string(),
locale: Locale::EnUs,
author: Author {
id: "78fc52a3-9f94-43c6-8eb5-9591e80b87e1".parse().unwrap(),
nickname: "HAL 9000".to_string(),
},
blog_id: "b4e05776-fca3-485e-be48-b1758cedd792".parse().unwrap(),
blog_title: "Binary encoding".to_string(),
};
let mut buffer = Vec::new();
SpecificSingleObjectWriter::<Comment>::with_capacity(10)
.unwrap()
.write_ref(&payload, &mut buffer)
.unwrap();
println!(
"Comment encoded with Single object encoding algorithm: {:02x?}",
buffer
); // c3 01 (Avro two byte marker) | fe c8 1c 7e df 23 cc dc (schema fingerprint) | 48 64 65 32 64 ... (payload)
To decode a message it is possible to use the local schema version or make a resolution by fetching the schema from the registry.
With local schema (not usable in production):
let comment = SpecificSingleObjectReader::<Comment>::new()
.unwrap()
.read(&mut buffer.as_slice())
.unwrap();
With schema resolution:
buffer.drain(0..2); // remove avro two-byte marker
let fingerprint: Vec<_> = buffer.drain(0..8).collect(); // extract schema fingerprint
let comment: Comment = from_value(
&from_avro_datum(
&fetch_schema(fingerprint.as_slice()).await, // fetch corresponding schema from the registry or local cache
&mut buffer.as_slice(),
Some(&Comment::get_schema()),
)
.unwrap(),
)
.unwrap();
Rust code is available on GitHub
Ecosystem
I had 0 knowledge of Avro before writing this article, below is my feedback on its ecosystem.
First, the documentation is not use case orientated but descriptive. Second, there are almost no examples/tutorials on the internet. Finally, the crate documentation, outside of the README, isn’t straightforward. However, the feature set looks complete according to the specification. But the API is not that user-friendly; the Rust implementation is a bit surprising sometimes.
Putting everything together as described above required many documentation search -> code read -> test cycles.
Everything is versioned in a mono repo at the time of writing. The specification is implemented in mainstream programming languages.
IMO, the available resources about Avro are too limited for easy adoption in a company. Moreover, the tooling is “old-fashioned” and of lower quality than mainstream technologies. Some efforts should be made on adoption by:
- Explaining the Avro protocol and its usage for this use case.
- Providing detailed examples with guidelines in all programming languages.
Next post presents how JSON, MessagePack and JSON schema can be used together to encode messages stored in an event bus.