0:00 Fergal O'Connor | KERIA & Signify Path to Production | KERI Conference 2026
0:03 My name is Fergal. I'm the
0:05 identity solutions lead architect at the
0:07 Cardano Foundation. I've been there for
0:09 the last four years building out our
0:11 open source identity solution known as
0:12 Veridian. Some of you may have seen
0:15 that this morning in Thomas'
0:16 presentation on the 5Ws.
0:18 Over a year ago, we launched the
0:22 Veridian wallet to the public app
0:25 stores. This is the first KERI/ACDC
0:28 mobile wallet in production. It's
0:30 based on KERI and Signify. We
0:33 invested a lot of time and effort into
0:36 bringing a nice user experience to
0:38 the user. You know abstracting away
0:41 the details of low-level KERI and
0:43 giving something that's actually usable
0:45 to the end user. So in doing that we
0:48 had to actually contribute a lot of
0:49 things back to KERI and Signify.
0:52 Yesterday I spoke about the security
0:55 architecture of KERI and Signify. But
0:58 today I'll talk about the path to
1:00 production. I think there are some
1:03 maturity gaps in the implementation of
1:05 KERI and Signify that should be addressed
1:08 for adoption of the KERI suite. It's
1:10 starting to lag behind the maturity of
1:11 KERIPy. We'll kind of step through
1:14 that today.
1:20 So, I mean yesterday's talk gave kind of an
1:23 in-depth explanation of how it works in
1:25 practice. But kind of the tagline
1:28 that I had was agent in the cloud
1:30 controller at the edge. So KERIA
1:33 allows you to deploy cloud agents
1:37 which essentially handle all of the
1:38 verification of KERI/ACDC/CESR up in the
1:42 cloud and allow you to use an edge
1:45 client library known as Signify at
1:47 the edge to just assign key
1:51 events essentially or ACDCs this
1:54 keeps the edge layer very thin.
1:56 For example, Veridian wallet on our mobile
1:59 just uses a Signify-ts to sign
2:02 things essentially and all of the
2:03 complex validation logic is pushed to
2:06 the cloud kind of like a black box
2:08 essentially. So in many ways it
2:10 becomes a more easy to use
2:14 API that's higher level
2:16 which is desirable for adoption
2:21 It gives something that's a bit
2:22 more tangible for people to develop on.
2:25 And very quickly, this is taken from
2:28 my deck yesterday, but doesn't look so
2:31 good on the four screens there but if
2:33 KERIA is this kind of block in the
2:35 middle is the multi- agent service it's
2:38 based on KERIPy. So it's a Python
2:40 service. It allows you to spin up
2:43 various agents in the cloud that are
2:45 fully isolated. So they have separate
2:46 databases. And for every agent in
2:49 the cloud, you have exactly one
2:50 Signify client at the edge. So this
2:51 these could be various applications from
2:54 various people sharing the same cloud
2:56 agent. This could be on a mobile
2:59 device for example. And there's a
3:02 secure communication between each client
3:04 and agent. Which I delved into
3:07 yesterday. I won't do that today.
3:10 And the agent is also responsible for
3:13 interfacing with everything else on the
3:15 network. Such as witnesses and
3:17 watchers and issuers and verifiers and
3:19 so forth. So Signify only ever talks
3:21 to the cloud agent.
3:23 – Quick question if I may. So why
3:26 must there only be one Signify client
3:30 per agent? So for example, if I
3:32 had a Veridian wallet and I had a
3:34 browser wallet. Why couldn't I both
3:37 connect to the same agent? What would be
3:39 the problem with that?
3:45 Do the Signify clients have different
3:47 key material?
3:48 – No, same password passcode.
3:51 Well, for me that's moving around the
3:52 passcode a bit too much.
3:54 Well, it's the same user though, right?
3:56 I have my phone and I have another two
3:58 phones or two devices. What's the
4:00 problem with that?
4:02 There wouldn't actually be a problem
4:04 with that if it's effectively the same
4:05 client. If you're using the same
4:06 passcode, it is effectively the same
4:08 client, just a different instance of it.
4:10 It'll derive the same key
4:12 material and authenticate in the same
4:13 way.
4:16 But if your applications keep local states
4:18 you might it might cause
4:21 inconsistencies essentially. So
4:22 – Okay. But it might work.
4:24 It would work. Yeah.
4:27 Although it might then .., how
4:31 it actually stores data on KERI needs
4:33 to be the same. You've seen this when
4:34 you've tried to connect your
4:37 wallet to what was our wallet. Yeah, it
4:40 works but you see some fuzzy stuff
4:41 that's related to our wallet only.
4:44 But yeah, this is generally the setup.
4:49 And today I've kind of split into
4:51 three different sections of what I will
4:53 cover on how we can mature KERI and
4:55 Signify. First is focused on the
4:58 developer experience. The second on
5:01 the operational resilience and third on
5:03 scalability. I think the first two
5:07 are the most important to do in the
5:09 short term. To really you know
5:12 increase the usability of the stack.
5:15 Scalability is more relevant in the long
5:18 term. A lot of people think about this
5:19 now but these first two points are more
5:22 relevant in my opinion and scalability
5:24 will be very relevant for SEDI though you
5:26 know if we're launching KERI
5:27 instances that want to serve thousands
5:29 and thousands and thousands of people
5:32 with wallets and scalability becomes a
5:34 big concern. And just to preface as
5:37 well, I mean this is going to be a talk
5:39 probably explaining all the bad points
5:40 of KERI and Signify, but it does work.
5:42 It is in production. It's just maybe
5:45 a bit nuanced to work with. It's very
5:48 unless you understand it, it can be hard
5:50 to work with because there's a lack of
5:51 feedback in places. But all of that is
5:54 really just traditional software
5:56 development points that are missing.
5:57 Just some extra maturity. It's not like
5:59 it's not anything related to the
6:00 protocols as such or the security.
6:03 And in places where I say that, oh,
6:06 and in places where I say that 'Oh,
6:07 there's not enough API hardening in
6:09 place', the reality of that is just
6:11 something happened. The KERIA crashes.
6:12 It's not a security bug. For as so much
6:15 as I can see, it's more about a
6:17 resilience thing; okay?
6:20 So, on the developer experience the
6:24 first point on this slide
6:26 refers to 'At the edge.' So everything in
6:29 Signify essentially.
6:34 We've been working quite hard over the
6:37 last at least 6 months I think to
6:40 strongly type the edge libraries so I
6:44 mean if you're using Typescript
6:48 you should be using types essentially
6:49 but when Signify was originally written
6:52 It didn't have any types in it
6:55 This is problematic if you're
6:56 refactoring code and so forth things
6:58 start to break. So we've put a lot
7:02 of time and effort to strongly type
7:04 Signify. It's not fully there yet,
7:06 but it pretty much almost is. And
7:09 instead of adding types directly into
7:11 Signify, we've actually taken a kind of
7:13 A different approach because as I
7:16 mentioned yesterday, Signify essentially
7:18 is an API wrapper around
7:20 KERIA. It can do extra things like
7:22 generate key events and sign them, but
7:24 it's otherwise an API wrapper. And
7:27 since you have an API on KERIA, you can
7:30 let KERIA decide what it wants and what
7:33 it will respond with through its open
7:35 API documentation. It will say
7:38 at this endpoint, it will give a
7:40 response in this format if it's a HTTP
7:43 200 or whatever. So,
7:46 and this is common practice, to take open
7:49 API documentation and autogenerate types
7:53 in languages. So that's what we've
7:55 done: We've updated all of the open API
7:58 docs on KERIA and generated types in
8:00 Signify-ts and Signify Java because they
8:03 propagate to various languages now as
8:06 a single source of truth and if you
8:07 wanted to add in another Signify like a
8:10 Rust library you could just
8:12 automatically generate them. Now you
8:16 might have to do some reworking in
8:17 places.
8:18 Java for example is
8:22 not as expressive in its type system as
8:24 TypeScript is. It didn't play nice
8:28 with a lot of the semantics of CESR.
8:30 But we've kind of got it working now.
8:32 So it took some changes there but
8:35 it's quite useful and also we're not
8:37 hand rolling the open API docs. We're
8:40 actually taking
8:42 the data structures in KERIPy that
8:44 Sam has written saying "Here are the
8:46 top level fields of an ACDC, we've
8:49 converted those into Python data classes
8:52 automatically generated so they can be
8:54 reused in KERIA and those data classes
8:56 are converted to the open API docs" and
9:00 you what that gives you is, if there's an
9:01 update you know downstream, it propagates
9:04 upstream and the open API docs don't go
9:06 out of sync. So this has been a huge
9:10 lift to make it more usable.
9:13 Internally, I think Signify still needs
9:15 some work.
9:18 Probably we can get away with not doing
9:20 that for a bit. There's maybe some other
9:21 points that are more relevant to that
9:23 point. I would say we'd maybe tackle
9:25 this when we're upgrading to CESR2,
9:27 which would be quite a bit of work as well.
9:29 So, quick question if I may. So
9:33 some of the WebOfTrust developers
9:35 refer to Signify meaning ..? Is it what
9:38 Signify-ts has in it? I always thought
9:40 of it as like the protocol that KERIA
9:43 exposes. So is Signify a component or is
9:47 it a protocol or is it all things or I
9:49 don't ... So how would you define what
9:51 Signify is? So sorry it's a simple
9:54 question, but it's a language question mostly.
9:57 An 'Edge client' that you can ...
10:01 – Any edge client of KERIA?
10:02 Yeah.
10:03 – Okay. Got it.
10:07 Yeah. There's a KERIA
10:09 protocol essentially,
10:10 – Right, it's defined as a protocol, but
10:12 it could be Signify JS or Signify Java
10:15 or
10:15 Yeah.
10:17 whatever other language.
10:18 How does it look like? Is it
10:21 UX stuff or is it command line stuff or
10:25 library?
10:26 Library. So when you're using Signify-ts,
10:29 it's like having
10:31 a compiler error if you type in the
10:32 wrong formats, essentially. So if you try
10:35 to create an ACDC or something, it will
10:38 know what the structure should look like
10:40 and you can reference it directly.
10:43 Whereas before you're kind of blind,
10:46 essentially.
10:49 Maybe you receive an ACDC back
10:51 from KERIA and you were going to present
10:52 that in the UI and you would say I want
10:55 to get the top level 'I' field as the
10:57 issuer.
10:59 If you were to reference some field that
11:01 doesn't exist, like 'T', before it
11:03 wouldn't complain and you just get a
11:04 blank space in the UI whereas now
11:07 TypeScript would say there is no top
11:10 level T field in an ACDC. So
11:14 I'm a strong believer in
11:16 strongly-typed languages.
11:19 The payoff is very high in terms of
11:21 keeping the bug count low and
11:23 especially when you refactor.
11:26 To your question: It's a library
11:28 you need to be able to build that you
11:30 want on top of
11:32 sort of a hole in the wall to service
11:44 food. Like you give to the other side.
11:47 The delivery desk.
11:50 Yeah,
11:50 you have the back office which is
11:55 KERIA
11:56 and then you have this (service window, red) and then you can
11:59 deliver it out to the UX (User Interface, red)
12:00 "Serving window" or something?
12:03 Yeah.
12:04 Okay. And with this you know what's
12:05 going to come out, what shape it will
12:07 be, what comes through the serving
12:09 window whereas before it's like you're
12:10 blindfolded.
12:14 No, I'm good. Another one which is
12:18 a bit technical. Currently with in
12:22 all of the Signify implementations
12:24 When you want to do something like
12:27 issue a credential
12:30 in one command on Signify
12:33 it will construct the credential
12:37 Try to create the interaction event
12:39 to anchor it in your Key Event Log. It
12:42 will pick the sequence number for you
12:44 and submit it to the API all in one go.
12:47 Absolutely.
12:48 Which automates workflows of several
12:50 steps. Yeah.
12:51 Yeah. Which is not a good thing because
12:53 if you lose the HTTP response
12:57 you've now lost data; okay? It's okay
12:59 when you use Signify completely
13:01 stateless in smaller applications but
13:02 when you want to build enterprise
13:05 applications where you need
13:07 mission critical reliability you need
13:10 things to go through. So if you want
13:11 outbox patterns, you need to be able to
13:13 decouple those things. So you can create
13:15 and sign at the edge and then submit.
13:19 And the API should be idempotent.
13:24 Yeah.
13:25 On that particular note, I want to
13:26 make I want to hear more what you have
13:28 to say on outbox patterns. That's
13:29 actually something I was looking into.
13:30 It seems like that while we're
13:32 considering this,
13:33 we should consider the API, the
13:36 fundamental design of the API between
13:38 Signify and KERIA because one of the
13:40 reasons why we have so many different
13:42 HTTP payloads instead of just CESR
13:44 streams which if we were to just had
13:46 streams the outbox it would be very
13:47 straightforward
13:48 is because we didn't have a fully a
13:51 complete parser and cryptographic
13:54 primitives implementation in TypeScript
13:56 when Signify-ts was being first created.
13:59 And so, we ended up doing, a bit of what
14:00 you would say is, a bulkier API
14:03 from Signify-ts to KERIA because
14:05 really we should just be passing CESR
14:07 strings around. And instead we have these
14:10 and embeds that we put in HTTP messages.
14:12 And so I it's sort of a something
14:15 that I've wondered about as part of
14:16 maturing the Signify and KERIA. But
14:19 even before we even get there, I do
14:22 think we should architect everything the
14:24 way that you're saying about outboxes:
14:26 save everything locally. That makes a
14:28 lot of sense. So
14:29 but I'm just sort of saying if we're
14:31 going to do that anyway, maybe we should
14:32 reconsider what we even send per
14:35 endpoint.
14:35 Yeah, that's a big refactor as well or
14:38 change rather. Yeah,
14:39 but you could put helper functions,
14:41 you know, also in Signify-ts, sorry,
14:44 that either give the structure or raw C.
14:48 Okay, going
14:48 I'll finish because I have like 10 sides
14:50 of this. This is on our fork. We've
14:54 done this already with identifier
14:55 creation. So, we create the identifier
14:56 locally and then submit it and it allows
14:58 us to get reliability in our wallet,
15:00 which we couldn't get before. But we
15:03 probably need to apply this everywhere.
15:06 Can you say that again? You create the
15:08 identifier and then submit it. Yeah. So
15:09 the post call is decoupled. So we
15:13 can put it into an outbox locally and
15:15 push it up. We've gone to a lot of
15:19 lengths for reliability in our wallet,
15:20 but you need to do things like that for
15:23 it. And you also get better control
15:25 in the sequence number for anchoring. So
15:27 if you're trying to assign multiple
15:28 things at once, especially in a group
15:29 multi context, having better control
15:32 will be better.
15:33 – How is this how is that just
15:35 different from Signify right now? Is that
15:37 what the 10 slides are
15:38 example, right now in Signify, your IPEX
15:42 is decoupled. You create the IPEX
15:43 message and you submit it.
15:45 It's two calls. Everything else isn't,
15:47 but it should all be like IPEX.
15:48 I see.
15:49 Yeah.
15:51 Wants to go through.
15:54 Maybe I should leave the room.
15:57 All right.
15:59 Okay. It's just an example of the
16:01 TypeScript generated,
16:03 a snippet of it. It's a very long file.
16:05 And in Java there's like hundreds of
16:07 classes generated but .. This is what
16:11 we're contributing. The other
16:14 developer experience issue, and this is
16:16 actually probably even a bigger issue,
16:19 And you see this all the time on
16:21 Discord for new developers is
16:23 Silent Failures
16:27 where you submit something to
16:29 KERIA
16:31 and nothing's happening essentially.
16:35 Now with KERI when you submit things
16:37 that take time you get an operation back
16:39 and you can track the operation
16:40 completion such as if it's waiting for
16:42 witness receipts and so forth. But
16:45 it's so common to have operations that
16:46 are stuck in pending forever and you
16:49 have no idea why and you can't trace it
16:51 down and it's because you had a
16:52 different SAID in this place of this .., and
16:54 there's very little feedback to why
16:57 that happened.
17:01 – You want to put more weight on it?
17:04 Yeah, you just need more you just need
17:06 better feedback patterns essentially.
17:07 It's probably a case by case basis
17:09 to do that. This is at kind of the
17:12 API layer. It needs it needs work.
17:14 – But there could be like a 'result key' kind
17:17 of pattern applied to the call so
17:20 that if there's a failure in any call,
17:23 it's handled in a consistent way/
17:25 Yeah. But this is not the call. The call
17:27 is like it accepts it and it puts it
17:29 onto a parser
17:31 and then the result of parsing that
17:34 doesn't result in a fully signed
17:35 something because something was wrong
17:37 and you have no idea why and you're
17:39 tracing through the logs and there's a
17:40 million logs.
17:44 I don't have a I mean this is case by
17:46 case basis so you'd have to step through
17:48 the code and find the biggest issues but
17:50 even if you historically look at the
17:51 GitHub issues or Discord you'll see what
17:53 pops up most. This one
17:58 I think we fixed this one and actually
17:59 caused a different issue or
18:02 not sure.
18:04 Yeah, we made a change. Yeah, cuz it
18:06 wasn't
18:08 anyway. Yeah, but in general it just
18:10 needs work. I would say it's a
18:11 barrier to entry. Yeah, I've seen
18:15 this too often where the operation just
18:16 times out and you have no idea why.
18:20 I would think they're probably the two
18:22 biggest things you could do for the
18:23 developer experience. So, it's actually
18:25 easy to build on, you know because
18:27 if you get stuck wondering why this
18:29 isn't working for hours and hours and
18:32 hours, it's
18:35 If you finally convince someone to try KERI
18:38 Yeah.
18:39 and they start developing and this is
18:40 experience.
18:42 Yeah. But this these are easy things to
18:45 fix. It just takes a little bit time to
18:46 go through certain edge cases. The
18:48 harder one to fix is the typing which
18:50 we've almost done. So, that's the good
18:51 news. The second so that was
18:54 developer experience. The second kind
18:56 of section is Operational Resilience
18:58 you know after you deploy it
19:00 what that looks like. Some of these
19:01 might bleed back into developer
19:03 experience though. Related to the
19:07 last one is API hardening. They're
19:11 kind of overlapping. But additionally
19:14 not just
19:16 you know KERIA getting stuck in the
19:17 state where things are pending. It's
19:19 also a bit crash prone from certain bad
19:21 inputs.
19:23 What happens if it crashes?
19:25 Well, if it's in the Docker setup, it
19:27 just comes back up. That's fine. But
19:28 we even see in our
19:32 developer KERI deployment that we use
19:35 at Veridian and you'll have some random
19:37 person in the community trying to do the
19:39 previous slide sending the wrong input
19:42 and they crash our instance and all of
19:44 our wallets are flashing to offline mode
19:46 because KERIA has gone down and up,
19:48 it needs a bit more
19:51 resilience, at that point. It's not
19:53 on the parser. It's actually just on
19:55 things like the granter and things like
19:57 that. All the dexter and the agent in
19:58 file.
19:59 It's the development, not the
20:00 production.
20:02 Oh, yeah. Yeah. Yeah.
20:03 He says our wallets are going up and
20:04 down. That's
20:05 the development environment.
20:06 Yeah. Yeah.
20:09 They don't have access to our
20:10 production.
20:14 And, of course, KERI is multi-
20:16 agent. Okay. So yeah, like I just said,
20:20 you're now denying service to other
20:22 people just because you're submitting
20:23 the wrong thing. And it bloats the
20:26 logs with lots of errors. That could be
20:28 nice,
20:29 you know, information back to the user.
20:31 You submitted the wrong thing.
20:36 Yeah, I believe
20:39 this was the one we fixed recently.
20:42 So it doesn't crash anymore if you
20:45 point your IPEX
20:48 Admit to the wrong ..., instead of
20:51 pointing it back to the grant, you point
20:53 it to a different string. It caused a
20:56 crash. Now it causes
20:59 a an operation that's pending forever.
21:01 So you just pushed it down one layer.
21:06 Another one is application layer
21:10 reliability.
21:12 I think there's kind of a misconception
21:17 that the escrows in KERIPy are
21:21 meant to provide application layer
21:22 reliability, but they're not. They're
21:24 more they're more so to
21:27 handle out-of-order processing; okay?
21:31 In reality, if you want full reliability
21:33 at the application layer between
21:36 two wallets, their agents will need to
21:39 continuously resend information
21:43 Until they get a acknowledgements
21:44 back. For example
21:48 currently if I send an IPEX grant and my
21:52 KERI sends it to the port of another
21:55 KERI instance and it's accepted 202
21:59 all that's done is:
22:00 put it onto the parser.
22:02 It's not going to give you any
22:04 feedback if it crashes. The parser is in
22:07 memory right? So if their KERI instance
22:10 goes down at that point in time when you
22:12 just got the 202 back, you have no idea
22:15 that they didn't parse it at all.
22:16 They've never seen it essentially. So
22:19 for application reliability, we need
22:21 retries on that.
22:23 – When you're talking about “they”, you
22:25 mean the other controllers agent
22:28 maybe not even related to your agent
22:30 other than it's on the same KERIA. So,
22:33 I was experiencing the need for
22:36 retries that I didn't understand because
22:37 I have a good network. But it
22:41 could just so as a defensive programming
22:45 when you get the 202 or other certain
22:47 other areas, you need to retry. It's not
22:49 necessarily a network error.
22:52 Yeah.
22:52 Is that what you're saying? It's the
22:53 KERIA is busy or dropped it.
22:56 Yeah. I mean, so on the 3902 port, the
23:00 CESR endpoint of KERIA, it will just
23:03 take in the CESR stream and put it
23:05 onto the queue of the parser and give
23:08 you a 202 back and then parse it. But if
23:12 it gives you the 202 back and
23:14 immediately crashes, you've now lost
23:16 that data.
23:18 They have no idea that they need to
23:19 parse something. They can't parse it.
23:21 It's lost in memory. But the 3902 port
23:23 isn't used by Signify-ts, right? It's
23:26 I'm not talking about Signify-ts here.
23:28 I'm talking about entire KERI here,
23:30 but agent-to-agent communication. Yeah.
23:35 but also from Signify clients.
23:38 can use it if they want to? 3902 is open.
23:43 Yeah. But Signify-ts doesn't do anything
23:45 with it, out-of-the box, right?
23:47 No, it doesn't out-of-the box. No.
23:50 But also, if your agent accepts,
23:53 an IPEX message, it should
23:56 send it, right?
24:00 Especially for groups. So, for example,
24:03 if you have a group set up
24:07 based on KERIA and Signify
24:10 Let's say, the one person proposes an
24:14 IPX message, sends it to the others, it
24:16 starts collecting signatures. before it
24:19 collects enough signatures, it goes
24:21 down, comes back up, receives the other
24:25 signatures, nothing happens. It doesn't
24:27 send it to the other person. Because the
24:29 granter deck is in memory, the queue.
24:32 So, it's lost it. And the lead of the
24:35 group currently is the one that sends.
24:37 So, a lot of the in-memory decks that we
24:39 use for cues in KERI need to be made
24:42 durable. So
24:45 Is there an issue? If there's not an
24:47 issue with that, I'm going to make that
24:48 one right now. That's
24:49 Yeah, that's bad. That one
24:51 that one's bad, actually. That needs to
24:52 be fixed soon. But you don't see
24:55 these in your normal setups like
24:57 because it's just "happy paths,"
25:00 everything just works. Everyone clicks
25:01 at the same time.
25:02 It's all good.
25:03 Yeah.
25:05 And it'd be nice to have some more
25:07 feedback on the async stuff because
25:08 everything in KERI is async. So,
25:11 if we have better ways of receiving
25:14 results that are more rich on what's
25:17 happening, that would be useful. But
25:20 that's a bit open-ended.
25:22 I mean, if you already have a plan to
25:23 put these into issues, I don't want to
25:24 duplicate it. But is there
25:27 a plan to currently make an
25:29 (GitHub, red) Issue on that because if there's not,
25:30 I'll capture it immediately.
25:31 I have a joint (GitHub, red) Discussion but I don't
25:34 know if it has everything from this. It
25:35 was written a while ago. It's either
25:37 in Signify-ts or KERIA (GitHub repos, red).
25:42 that's some information from Sam on application layer reliability.
25:48 If you want to read it, it's him saying
25:49 that KERIPy is not going to provide
25:52 this essentially.
26:01 What is him saying again? About achieving application layer
26:04 reliability end to end.
26:06 Okay, cool. And the need for retries and
26:08 idempotency, so forth.
26:15 Cool. Finally, well, not finally.
26:25 Especially when you're using group
26:27 multisig, and if you're doing something
26:29 wrong, it can be quite hard to
26:31 understand what you're doing wrong
26:33 because it's a multi-agent setup. All
26:36 your logs are intertwined. There's no
26:38 way to filter on the logs. And even if
26:40 there was a way to filter on the logs
26:41 per agent, there's still too many
26:44 logs, there's no way to filter on the
26:45 action you were trying to do. And
26:49 data again, everything's based on HIO.
26:52 Everything's async, everything's moving
26:54 various ways with little
26:57 observability. It's again fine,
27:00 once you've developed it's fine,
27:01 but sometimes it can be very
27:04 difficult to add new features to KERIA
27:05 and try and figure out where the data is
27:07 actually moving and what's actually
27:09 wrong and why it isn't completing and
27:11 so forth.
27:12 Should the KERIA requests have a
27:14 request ID added to them?
27:15 Probably. Yeah. But it'll need
27:18 some work with KERIPy as well
27:21 because most of the complex logic of
27:24 KERIA is KERIpy. It's a direct
27:26 dependency. It gets pushed onto KERIPy.
27:28 So I need to do some work
27:30 there. So yeah, some instrumentation
27:34 would be useful at that.
27:37 You see you go through the
27:39 logs and you're just trying to figure
27:40 out what's going on, because you've put
27:41 the wrong ID somewhere, and it's ...,
27:44 You can see Kent added some
27:46 stuff here. But we probably need
27:48 to go a bit further so we can actually
27:50 filter on IDs of agents and so forth.
27:53 A structured log.
27:54 Yeah.
27:55 Request ID or client ID.
27:56 Yeah.
27:59 I'll do this one quite quickly.
28:04 Yeah, probably KERI and Signify should
28:05 have some semantic versioning because
28:07 it's an API. And I think you raised
28:10 issues on this. There should be
28:13 compatibility enforcements between them.
28:15 I don't think this is big issue now,
28:17 but in as you move towards
28:20 being a more mature ecosystem and
28:21 getting more developers in, it'll become
28:23 more relevant. Also versioning.
28:28 Right now you have to version on the
28:30 KLI per agent and iterate over them. So
28:34 there needs to be a migration framework
28:36 in place. We've been avoiding this
28:38 for ages because there hasn't been
28:39 migrations, essentially, but when there is
28:42 in production it would be
28:43 problematic. So that's resilience.
28:49 And again these are
28:54 traditional software problems
28:57 Just put some resources in
28:58 and it'll make it a more mature app
29:00 It's not a protocol thing. These are
29:02 all very achievable.
29:05 Finally on Scalability
29:08 And this one is also a kind of
29:11 developer experience
29:14 I mentioned earlier that because a lot
29:16 of the things are async, you need to
29:19 submit like create an identifier, create
29:22 a registry, issue a credential and you
29:23 get an operation back.
29:26 And then you need to pull on that
29:28 status of the operation until it's done
29:30 while it's waiting to collect signatures
29:31 and so forth. That's kind of a pain.
29:35 Typical modern systems are usually
29:38 event driven. Where you would get
29:40 some notification that it has completed
29:42 rather than constantly checking.
29:46 The same goes for notifications. You
29:47 have to constantly check if there's new
29:49 new notifications.
29:52 You know in any general application like
29:54 a wallet, it's going to have to
29:55 constantly check because it doesn't know
29:57 when to expect that you're receiving a
29:59 credential.
30:00 – Single-threaded environment, it could
30:01 completely
30:04 block the rest of the application.
30:05 Yeah.
30:08 We've made event driven in our
30:10 wallet in every application we've built
30:11 because we've just put it in a corner in
30:13 the box and made it event driven from
30:16 the perspective of the rest of our
30:17 application. But it's kind of annoying
30:19 to have to build that every single time.
30:20 But this is very doable, if we were to
30:24 take the query reply
30:26 approach of just using the
30:29 the WSGI response stream, we could open
30:32 up the server ... connection per
30:34 agent. Yep.
30:36 To Signify and do this today.
30:39 I might try that over the weekend.
30:43 Huh?
30:45 No.
30:46 Tell me about it.
30:52 Before I get to that. The reason I put this in scalability is
30:54 because right now you have 10,000
30:56 wallets using one KERI nstance and
30:59 they have 25 pending operations that
31:01 they're pulling, pulling, pulling. These
31:03 are all full HTTP requests. It's way too
31:06 much traffic hitting KERIA.
31:11 It won't work. So moving to something
31:15 like Server Sent Events
31:17 is much better from a developer
31:18 experience, but it's also much better
31:20 from a load experience. And
31:24 it isn't, ..., I've never
31:26 implemented them, so I don't know what
31:27 it's like, but for notifications, it
31:30 wouldn't be too difficult because it's
31:32 just the next notification coming out on
31:34 the event. It's more problematic for
31:36 operations. I explain that in a moment.
31:39 I also Server Sent Events are true
31:42 HTTP, so they're compatible with load
31:44 balancers. So if you start to do
31:46 horizontal scaling and so forth, they're
31:47 compatible, which is why Sam
31:50 would prefer that over
31:51 websockets because that isn't pure HTTP.
31:59 Now the problem for integrating
32:01 with operations is currently
32:02 when you perform an action that requires
32:07 a long run operation to track that
32:09 state of the operation.
32:13 It just stores the related SAIDs and
32:16 metadata that refer to that operation.
32:18 It doesn't have a status of pending or
32:20 failed or anything like that.
32:22 And when you use Signify to check that
32:24 status of pending or failed, it's not
32:26 stored in the database. That's what I
32:27 mean. It doesn't flip in the database
32:29 from pending to complete
32:31 actually if you look at longridden.py
32:35 in KERI, it takes all of those SAIDs
32:38 and compares it against the database of
32:40 the agent to see if it's effectively
32:42 complete. Okay, so it's not checking is
32:46 this operation complete in the database.
32:48 check in based on this operations
32:50 metadata are all the signatures in place
32:53 for this to be considered complete for
32:55 example you can see here for the
32:57 delegations using the 'swain.complete'
33:00 every time you check
33:03 Yeah, basically like a diff on the
33:05 database
33:06 – Yeah, this has been quite a mental
33:08 challenge for me for like two years with
33:10 KERI is the notion that the concept of
33:13 state and state modeling is missing
33:16 and that's intentional
33:18 I think from Sam
33:22 IPEX for example, right, is ..
33:25 I tried to get a state
33:28 machine diagram of all the possible
33:31 state transitions and IPEX of Spurn,
33:33 Grant, Accept, and all that and
33:36 I had a challenge getting to the
33:38 precision but this is similar: You're
33:41 wanting to query the state of an
33:44 operation and it has to be derived and
33:46 has to be requeried to compute the state.
33:48 That is what I think I heard you say.
33:51 Yeah.
33:51 ...of an operation, right? So it'd be
33:54 nice just to get the state back
33:56 from KERIA like 'Oh, it's in Pending'.
33:58 Yeah. So basically while currently
34:00 we can't do event driven with
34:01 Server Sent Events.
34:03 To date?
34:04 Because it's you're checking against the
34:06 database state. Okay. So instead of
34:09 checking the database state or the diff
34:13 when this completes for the first time
34:15 inside KERIPy, it needs to emit an
34:17 event that gets sent on server the
34:19 Server Sent Event
34:21 thread.
34:22 Okay. Because I'm thinking what we could
34:23 do for each operation type is that there
34:27 could be there's already a
34:28 server event connection from the agent
34:30 to the Signify client. It would
34:32 essentially just be a channel to
34:33 communicate. I'm thinking on completion
34:35 it would just send a completion.
34:36 Yeah. But there is no completion of an
34:39 operation in KERIA. It's a state check.
34:42 It tells you it's effectively complete.
34:46 There's no point where it's in the
34:47 database ..,
34:48 – Iit's inferred
34:49 Yeah, it's inferred. So if you
34:51 never ever connect it again,
34:53 nothing changes on the database to say
34:55 it's complete. It's just you coming back
34:57 later and saying this is the status.
35:00 Okay. So because there's no point at
35:02 which that happens
35:04 This is a problem all the
35:06 time, right?
35:06 So essentially
35:08 we need to store the completed state, not
35:10 just derive it.
35:11 Yeah, All of this happens
35:13 in KERIPy so KERIPy needs to tell us
35:16 because this is a KERIA thing
35:18 operations aren't in KERIPy
35:20 – .. locally in KERIA per agent so because
35:22 once we derive it once even though we're
35:24 reading it
35:24 yeah but what triggers to derive it
35:26 other than you're going to move the
35:28 polling inside internally
35:29 – What triggers what?
35:31 What triggers 'Check and it's complete.'
35:34 in KERIA.
35:35 So this is where I think what we should
35:37 do when we have an operation
35:39 There should be a Doer for each
35:42 operation
35:43 once it completes at the end before that
35:46 Doer finishes it should write the
35:47 complete state.
35:48 Yeah. But now you've just
35:50 moved all of the polling internal to KERIA
35:53 and you'd have so many doers for 10,000
35:57 Signified clients. So it would
36:00 be a lot better if internally in KERIPy
36:03 It could just let us know when it's
36:05 effectively completed
36:08 because there's already something inside
36:09 KERIPy tracking this I would say. So
36:11 then it can admit that event or whatever
36:14 and we can capture it and update our
36:17 operation in the database and send the
36:19 event back and Server Sent Events.
36:21 Yeah, I think we're slightly talking
36:22 past each other but I think I get what
36:23 you're saying. I think fundamentally
36:25 when an important event happens in KERIPy,
36:27 it should be admitted saying "this and this
36:29 happened."
36:30 Yeah.
36:32 Rather than having new Doers that are
36:33 also tracking it again.
36:37 Cool, just one more slide. So it's
36:38 actually good.
36:39 No, actually one; so not too bad.
36:44 Yeah, and horizontal scaling.
36:46 People have asked about this before
36:49 I would put this on the lowest priority.
36:50 because other things just need to be
36:52 fixed first to get adoption. But for
36:54 SEDI it would be important.
36:57 Okay, right now you could do something
36:59 like sticky sessions. It would be a very
37:00 naive approach. Where you
37:04 deploy 10 different KERIAs and you
37:06 have a load balancer and once a
37:10 client connects and gets routed to a
37:12 KERIA instance they always have to go
37:13 there, forever. That's like a sticky
37:15 session. It would work now, but it
37:17 wouldn't be great.
37:20 To do true horizontal scale, I think KERI is
37:23 actually already quite suited to it.
37:26 Because all the databases of agents are
37:27 fully isolated and no agents are shared.
37:31 They're just a bunch of LMDB files.
37:34 So if you had to ... [reformulates, red.]
37:40 The easiest thing you could do is have a
37:42 shared volume in your cloud deployment
37:43 that all of the instances can access and
37:45 they would just go to the specific LMDB
37:47 files that they're currently
37:49 serving. And if it's done correctly
37:52 they won't try to access the same LMDB
37:53 file of someone else and you could have
37:55 the agency coordinate that. If you
37:58 wanted more isolation, I guess you could
38:01 separate the volumes, but I think that
38:02 would be
38:05 a little bit complex, not necessary.
38:11 You'd be moving data across
38:13 volumes. I don't like that for now.
38:14 At least I don't think it's necessary.
38:16 Unless there's a
38:20 legal constraint on data or something.
38:24 And yeah, the agency can become the
38:26 load balancer, like it currently is
38:28 basically for the message router in
38:29 KERIA. But instead it would
38:31 actually be outside
38:33 KERIA and route into a more
38:38 light-weight KERIA. So you kind of
38:40 decouple it a bit; be different services.
38:45 I think that's more down the line.
38:46 So that's why developer experience
38:48 and operational resilience are highest
38:50 priority, I think. Yeah, Done!
38:59 – What's this up there?
39:00 That's the file system.
39:03 The agency database is an LMDB
39:05 database. You can see that there was two
39:08 new agents booted. There are the config
39:12 files that get copied in the JSON config
39:15 files with the OOBIs.
39:17 And for each
39:20 agent they have a subdirectory in all
39:23 of these directories. Like there's the
39:25 key store, for example.
39:29 Probably it would be better if you
39:31 restructure it, so that the agent is at
39:34 the top level and it has its databases
39:36 internally so you can more easily take
39:38 it. But if it was a shared volume
39:40 across many instances it wouldn't
39:43 matter.