KERIA & Signify Path to Production - Fergal O'Connor

KERICONF26 Day 2 · 38:45

0:00 Fergal O'Connor | KERIA & Signify Path to Production | KERI Conference 2026

0:03 My name is Fergal. I'm the

0:05 identity solutions lead architect at the

0:07 Cardano Foundation. I've been there for

0:09 the last four years building out our

0:11 open source identity solution known as

0:12 Veridian. Some of you may have seen

0:15 that this morning in Thomas'

0:16 presentation on the 5Ws.

0:18 Over a year ago, we launched the

0:22 Veridian wallet to the public app

0:25 stores. This is the first KERI/ACDC

0:28 mobile wallet in production. It's

0:30 based on KERI and Signify. We

0:33 invested a lot of time and effort into

0:36 bringing a nice user experience to

0:38 the user. You know abstracting away

0:41 the details of low-level KERI and

0:43 giving something that's actually usable

0:45 to the end user. So in doing that we

0:48 had to actually contribute a lot of

0:49 things back to KERI and Signify.

0:52 Yesterday I spoke about the security

0:55 architecture of KERI and Signify. But

0:58 today I'll talk about the path to

1:00 production. I think there are some

1:03 maturity gaps in the implementation of

1:05 KERI and Signify that should be addressed

1:08 for adoption of the KERI suite. It's

1:10 starting to lag behind the maturity of

1:11 KERIPy. We'll kind of step through

1:14 that today.

1:20 So, I mean yesterday's talk gave kind of an

1:23 in-depth explanation of how it works in

1:25 practice. But kind of the tagline

1:28 that I had was agent in the cloud

1:30 controller at the edge. So KERIA

1:33 allows you to deploy cloud agents

1:37 which essentially handle all of the

1:38 verification of KERI/ACDC/CESR up in the

1:42 cloud and allow you to use an edge

1:45 client library known as Signify at

1:47 the edge to just assign key

1:51 events essentially or ACDCs this

1:54 keeps the edge layer very thin.

1:56 For example, Veridian wallet on our mobile

1:59 just uses a Signify-ts to sign

2:02 things essentially and all of the

2:03 complex validation logic is pushed to

2:06 the cloud kind of like a black box

2:08 essentially. So in many ways it

2:10 becomes a more easy to use

2:14 API that's higher level

2:16 which is desirable for adoption

2:21 It gives something that's a bit

2:22 more tangible for people to develop on.

2:25 And very quickly, this is taken from

2:28 my deck yesterday, but doesn't look so

2:31 good on the four screens there but if

2:33 KERIA is this kind of block in the

2:35 middle is the multi- agent service it's

2:38 based on KERIPy. So it's a Python

2:40 service. It allows you to spin up

2:43 various agents in the cloud that are

2:45 fully isolated. So they have separate

2:46 databases. And for every agent in

2:49 the cloud, you have exactly one

2:50 Signify client at the edge. So this

2:51 these could be various applications from

2:54 various people sharing the same cloud

2:56 agent. This could be on a mobile

2:59 device for example. And there's a

3:02 secure communication between each client

3:04 and agent. Which I delved into

3:07 yesterday. I won't do that today.

3:10 And the agent is also responsible for

3:13 interfacing with everything else on the

3:15 network. Such as witnesses and

3:17 watchers and issuers and verifiers and

3:19 so forth. So Signify only ever talks

3:21 to the cloud agent.

3:23 – Quick question if I may. So why

3:26 must there only be one Signify client

3:30 per agent? So for example, if I

3:32 had a Veridian wallet and I had a

3:34 browser wallet. Why couldn't I both

3:37 connect to the same agent? What would be

3:39 the problem with that?

3:45 Do the Signify clients have different

3:47 key material?

3:48 – No, same password passcode.

3:51 Well, for me that's moving around the

3:52 passcode a bit too much.

3:54 Well, it's the same user though, right?

3:56 I have my phone and I have another two

3:58 phones or two devices. What's the

4:00 problem with that?

4:02 There wouldn't actually be a problem

4:04 with that if it's effectively the same

4:05 client. If you're using the same

4:06 passcode, it is effectively the same

4:08 client, just a different instance of it.

4:10 It'll derive the same key

4:12 material and authenticate in the same

4:13 way.

4:16 But if your applications keep local states

4:18 you might it might cause

4:21 inconsistencies essentially. So

4:22 – Okay. But it might work.

4:24 It would work. Yeah.

4:27 Although it might then .., how

4:31 it actually stores data on KERI needs

4:33 to be the same. You've seen this when

4:34 you've tried to connect your

4:37 wallet to what was our wallet. Yeah, it

4:40 works but you see some fuzzy stuff

4:41 that's related to our wallet only.

4:44 But yeah, this is generally the setup.

4:49 And today I've kind of split into

4:51 three different sections of what I will

4:53 cover on how we can mature KERI and

4:55 Signify. First is focused on the

4:58 developer experience. The second on

5:01 the operational resilience and third on

5:03 scalability. I think the first two

5:07 are the most important to do in the

5:09 short term. To really you know

5:12 increase the usability of the stack.

5:15 Scalability is more relevant in the long

5:18 term. A lot of people think about this

5:19 now but these first two points are more

5:22 relevant in my opinion and scalability

5:24 will be very relevant for SEDI though you

5:26 know if we're launching KERI

5:27 instances that want to serve thousands

5:29 and thousands and thousands of people

5:32 with wallets and scalability becomes a

5:34 big concern. And just to preface as

5:37 well, I mean this is going to be a talk

5:39 probably explaining all the bad points

5:40 of KERI and Signify, but it does work.

5:42 It is in production. It's just maybe

5:45 a bit nuanced to work with. It's very

5:48 unless you understand it, it can be hard

5:50 to work with because there's a lack of

5:51 feedback in places. But all of that is

5:54 really just traditional software

5:56 development points that are missing.

5:57 Just some extra maturity. It's not like

5:59 it's not anything related to the

6:00 protocols as such or the security.

6:03 And in places where I say that, oh,

6:06 and in places where I say that 'Oh,

6:07 there's not enough API hardening in

6:09 place', the reality of that is just

6:11 something happened. The KERIA crashes.

6:12 It's not a security bug. For as so much

6:15 as I can see, it's more about a

6:17 resilience thing; okay?

6:20 So, on the developer experience the

6:24 first point on this slide

6:26 refers to 'At the edge.' So everything in

6:29 Signify essentially.

6:34 We've been working quite hard over the

6:37 last at least 6 months I think to

6:40 strongly type the edge libraries so I

6:44 mean if you're using Typescript

6:48 you should be using types essentially

6:49 but when Signify was originally written

6:52 It didn't have any types in it

6:55 This is problematic if you're

6:56 refactoring code and so forth things

6:58 start to break. So we've put a lot

7:02 of time and effort to strongly type

7:04 Signify. It's not fully there yet,

7:06 but it pretty much almost is. And

7:09 instead of adding types directly into

7:11 Signify, we've actually taken a kind of

7:13 A different approach because as I

7:16 mentioned yesterday, Signify essentially

7:18 is an API wrapper around

7:20 KERIA. It can do extra things like

7:22 generate key events and sign them, but

7:24 it's otherwise an API wrapper. And

7:27 since you have an API on KERIA, you can

7:30 let KERIA decide what it wants and what

7:33 it will respond with through its open

7:35 API documentation. It will say

7:38 at this endpoint, it will give a

7:40 response in this format if it's a HTTP

7:43 200 or whatever. So,

7:46 and this is common practice, to take open

7:49 API documentation and autogenerate types

7:53 in languages. So that's what we've

7:55 done: We've updated all of the open API

7:58 docs on KERIA and generated types in

8:00 Signify-ts and Signify Java because they

8:03 propagate to various languages now as

8:06 a single source of truth and if you

8:07 wanted to add in another Signify like a

8:10 Rust library you could just

8:12 automatically generate them. Now you

8:16 might have to do some reworking in

8:17 places.

8:18 Java for example is

8:22 not as expressive in its type system as

8:24 TypeScript is. It didn't play nice

8:28 with a lot of the semantics of CESR.

8:30 But we've kind of got it working now.

8:32 So it took some changes there but

8:35 it's quite useful and also we're not

8:37 hand rolling the open API docs. We're

8:40 actually taking

8:42 the data structures in KERIPy that

8:44 Sam has written saying "Here are the

8:46 top level fields of an ACDC, we've

8:49 converted those into Python data classes

8:52 automatically generated so they can be

8:54 reused in KERIA and those data classes

8:56 are converted to the open API docs" and

9:00 you what that gives you is, if there's an

9:01 update you know downstream, it propagates

9:04 upstream and the open API docs don't go

9:06 out of sync. So this has been a huge

9:10 lift to make it more usable.

9:13 Internally, I think Signify still needs

9:15 some work.

9:18 Probably we can get away with not doing

9:20 that for a bit. There's maybe some other

9:21 points that are more relevant to that

9:23 point. I would say we'd maybe tackle

9:25 this when we're upgrading to CESR2,

9:27 which would be quite a bit of work as well.

9:29 So, quick question if I may. So

9:33 some of the WebOfTrust developers

9:35 refer to Signify meaning ..? Is it what

9:38 Signify-ts has in it? I always thought

9:40 of it as like the protocol that KERIA

9:43 exposes. So is Signify a component or is

9:47 it a protocol or is it all things or I

9:49 don't ... So how would you define what

9:51 Signify is? So sorry it's a simple

9:54 question, but it's a language question mostly.

9:57 An 'Edge client' that you can ...

10:01 – Any edge client of KERIA?

10:02 Yeah.

10:03 – Okay. Got it.

10:07 Yeah. There's a KERIA

10:09 protocol essentially,

10:10 – Right, it's defined as a protocol, but

10:12 it could be Signify JS or Signify Java

10:15 or

10:15 Yeah.

10:17 whatever other language.

10:18 How does it look like? Is it

10:21 UX stuff or is it command line stuff or

10:25 library?

10:26 Library. So when you're using Signify-ts,

10:29 it's like having

10:31 a compiler error if you type in the

10:32 wrong formats, essentially. So if you try

10:35 to create an ACDC or something, it will

10:38 know what the structure should look like

10:40 and you can reference it directly.

10:43 Whereas before you're kind of blind,

10:46 essentially.

10:49 Maybe you receive an ACDC back

10:51 from KERIA and you were going to present

10:52 that in the UI and you would say I want

10:55 to get the top level 'I' field as the

10:57 issuer.

10:59 If you were to reference some field that

11:01 doesn't exist, like 'T', before it

11:03 wouldn't complain and you just get a

11:04 blank space in the UI whereas now

11:07 TypeScript would say there is no top

11:10 level T field in an ACDC. So

11:14 I'm a strong believer in

11:16 strongly-typed languages.

11:19 The payoff is very high in terms of

11:21 keeping the bug count low and

11:23 especially when you refactor.

11:26 To your question: It's a library

11:28 you need to be able to build that you

11:30 want on top of

11:32 sort of a hole in the wall to service

11:44 food. Like you give to the other side.

11:47 The delivery desk.

11:50 Yeah,

11:50 you have the back office which is

11:55 KERIA

11:56 and then you have this (service window, red) and then you can

11:59 deliver it out to the UX (User Interface, red)

12:00 "Serving window" or something?

12:03 Yeah.

12:04 Okay. And with this you know what's

12:05 going to come out, what shape it will

12:07 be, what comes through the serving

12:09 window whereas before it's like you're

12:10 blindfolded.

12:14 No, I'm good. Another one which is

12:18 a bit technical. Currently with in

12:22 all of the Signify implementations

12:24 When you want to do something like

12:27 issue a credential

12:30 in one command on Signify

12:33 it will construct the credential

12:37 Try to create the interaction event

12:39 to anchor it in your Key Event Log. It

12:42 will pick the sequence number for you

12:44 and submit it to the API all in one go.

12:47 Absolutely.

12:48 Which automates workflows of several

12:50 steps. Yeah.

12:51 Yeah. Which is not a good thing because

12:53 if you lose the HTTP response

12:57 you've now lost data; okay? It's okay

12:59 when you use Signify completely

13:01 stateless in smaller applications but

13:02 when you want to build enterprise

13:05 applications where you need

13:07 mission critical reliability you need

13:10 things to go through. So if you want

13:11 outbox patterns, you need to be able to

13:13 decouple those things. So you can create

13:15 and sign at the edge and then submit.

13:19 And the API should be idempotent.

13:24 Yeah.

13:25 On that particular note, I want to

13:26 make I want to hear more what you have

13:28 to say on outbox patterns. That's

13:29 actually something I was looking into.

13:30 It seems like that while we're

13:32 considering this,

13:33 we should consider the API, the

13:36 fundamental design of the API between

13:38 Signify and KERIA because one of the

13:40 reasons why we have so many different

13:42 HTTP payloads instead of just CESR

13:44 streams which if we were to just had

13:46 streams the outbox it would be very

13:47 straightforward

13:48 is because we didn't have a fully a

13:51 complete parser and cryptographic

13:54 primitives implementation in TypeScript

13:56 when Signify-ts was being first created.

13:59 And so, we ended up doing, a bit of what

14:00 you would say is, a bulkier API

14:03 from Signify-ts to KERIA because

14:05 really we should just be passing CESR

14:07 strings around. And instead we have these

14:10 and embeds that we put in HTTP messages.

14:12 And so I it's sort of a something

14:15 that I've wondered about as part of

14:16 maturing the Signify and KERIA. But

14:19 even before we even get there, I do

14:22 think we should architect everything the

14:24 way that you're saying about outboxes:

14:26 save everything locally. That makes a

14:28 lot of sense. So

14:29 but I'm just sort of saying if we're

14:31 going to do that anyway, maybe we should

14:32 reconsider what we even send per

14:35 endpoint.

14:35 Yeah, that's a big refactor as well or

14:38 change rather. Yeah,

14:39 but you could put helper functions,

14:41 you know, also in Signify-ts, sorry,

14:44 that either give the structure or raw C.

14:48 Okay, going

14:48 I'll finish because I have like 10 sides

14:50 of this. This is on our fork. We've

14:54 done this already with identifier

14:55 creation. So, we create the identifier

14:56 locally and then submit it and it allows

14:58 us to get reliability in our wallet,

15:00 which we couldn't get before. But we

15:03 probably need to apply this everywhere.

15:06 Can you say that again? You create the

15:08 identifier and then submit it. Yeah. So

15:09 the post call is decoupled. So we

15:13 can put it into an outbox locally and

15:15 push it up. We've gone to a lot of

15:19 lengths for reliability in our wallet,

15:20 but you need to do things like that for

15:23 it. And you also get better control

15:25 in the sequence number for anchoring. So

15:27 if you're trying to assign multiple

15:28 things at once, especially in a group

15:29 multi context, having better control

15:32 will be better.

15:33 – How is this how is that just

15:35 different from Signify right now? Is that

15:37 what the 10 slides are

15:38 example, right now in Signify, your IPEX

15:42 is decoupled. You create the IPEX

15:43 message and you submit it.

15:45 It's two calls. Everything else isn't,

15:47 but it should all be like IPEX.

15:48 I see.

15:49 Yeah.

15:51 Wants to go through.

15:54 Maybe I should leave the room.

15:57 All right.

15:59 Okay. It's just an example of the

16:01 TypeScript generated,

16:03 a snippet of it. It's a very long file.

16:05 And in Java there's like hundreds of

16:07 classes generated but .. This is what

16:11 we're contributing. The other

16:14 developer experience issue, and this is

16:16 actually probably even a bigger issue,

16:19 And you see this all the time on

16:21 Discord for new developers is

16:23 Silent Failures

16:27 where you submit something to

16:29 KERIA

16:31 and nothing's happening essentially.

16:35 Now with KERI when you submit things

16:37 that take time you get an operation back

16:39 and you can track the operation

16:40 completion such as if it's waiting for

16:42 witness receipts and so forth. But

16:45 it's so common to have operations that

16:46 are stuck in pending forever and you

16:49 have no idea why and you can't trace it

16:51 down and it's because you had a

16:52 different SAID in this place of this .., and

16:54 there's very little feedback to why

16:57 that happened.

17:01 – You want to put more weight on it?

17:04 Yeah, you just need more you just need

17:06 better feedback patterns essentially.

17:07 It's probably a case by case basis

17:09 to do that. This is at kind of the

17:12 API layer. It needs it needs work.

17:14 – But there could be like a 'result key' kind

17:17 of pattern applied to the call so

17:20 that if there's a failure in any call,

17:23 it's handled in a consistent way/

17:25 Yeah. But this is not the call. The call

17:27 is like it accepts it and it puts it

17:29 onto a parser

17:31 and then the result of parsing that

17:34 doesn't result in a fully signed

17:35 something because something was wrong

17:37 and you have no idea why and you're

17:39 tracing through the logs and there's a

17:40 million logs.

17:44 I don't have a I mean this is case by

17:46 case basis so you'd have to step through

17:48 the code and find the biggest issues but

17:50 even if you historically look at the

17:51 GitHub issues or Discord you'll see what

17:53 pops up most. This one

17:58 I think we fixed this one and actually

17:59 caused a different issue or

18:02 not sure.

18:04 Yeah, we made a change. Yeah, cuz it

18:06 wasn't

18:08 anyway. Yeah, but in general it just

18:10 needs work. I would say it's a

18:11 barrier to entry. Yeah, I've seen

18:15 this too often where the operation just

18:16 times out and you have no idea why.

18:20 I would think they're probably the two

18:22 biggest things you could do for the

18:23 developer experience. So, it's actually

18:25 easy to build on, you know because

18:27 if you get stuck wondering why this

18:29 isn't working for hours and hours and

18:32 hours, it's

18:35 If you finally convince someone to try KERI

18:38 Yeah.

18:39 and they start developing and this is

18:40 experience.

18:42 Yeah. But this these are easy things to

18:45 fix. It just takes a little bit time to

18:46 go through certain edge cases. The

18:48 harder one to fix is the typing which

18:50 we've almost done. So, that's the good

18:51 news. The second so that was

18:54 developer experience. The second kind

18:56 of section is Operational Resilience

18:58 you know after you deploy it

19:00 what that looks like. Some of these

19:01 might bleed back into developer

19:03 experience though. Related to the

19:07 last one is API hardening. They're

19:11 kind of overlapping. But additionally

19:14 not just

19:16 you know KERIA getting stuck in the

19:17 state where things are pending. It's

19:19 also a bit crash prone from certain bad

19:21 inputs.

19:23 What happens if it crashes?

19:25 Well, if it's in the Docker setup, it

19:27 just comes back up. That's fine. But

19:28 we even see in our

19:32 developer KERI deployment that we use

19:35 at Veridian and you'll have some random

19:37 person in the community trying to do the

19:39 previous slide sending the wrong input

19:42 and they crash our instance and all of

19:44 our wallets are flashing to offline mode

19:46 because KERIA has gone down and up,

19:48 it needs a bit more

19:51 resilience, at that point. It's not

19:53 on the parser. It's actually just on

19:55 things like the granter and things like

19:57 that. All the dexter and the agent in

19:58 file.

19:59 It's the development, not the

20:00 production.

20:02 Oh, yeah. Yeah. Yeah.

20:03 He says our wallets are going up and

20:04 down. That's

20:05 the development environment.

20:06 Yeah. Yeah.

20:09 They don't have access to our

20:10 production.

20:14 And, of course, KERI is multi-

20:16 agent. Okay. So yeah, like I just said,

20:20 you're now denying service to other

20:22 people just because you're submitting

20:23 the wrong thing. And it bloats the

20:26 logs with lots of errors. That could be

20:28 nice,

20:29 you know, information back to the user.

20:31 You submitted the wrong thing.

20:36 Yeah, I believe

20:39 this was the one we fixed recently.

20:42 So it doesn't crash anymore if you

20:45 point your IPEX

20:48 Admit to the wrong ..., instead of

20:51 pointing it back to the grant, you point

20:53 it to a different string. It caused a

20:56 crash. Now it causes

20:59 a an operation that's pending forever.

21:01 So you just pushed it down one layer.

21:06 Another one is application layer

21:10 reliability.

21:12 I think there's kind of a misconception

21:17 that the escrows in KERIPy are

21:21 meant to provide application layer

21:22 reliability, but they're not. They're

21:24 more they're more so to

21:27 handle out-of-order processing; okay?

21:31 In reality, if you want full reliability

21:33 at the application layer between

21:36 two wallets, their agents will need to

21:39 continuously resend information

21:43 Until they get a acknowledgements

21:44 back. For example

21:48 currently if I send an IPEX grant and my

21:52 KERI sends it to the port of another

21:55 KERI instance and it's accepted 202

21:59 all that's done is:

22:00 put it onto the parser.

22:02 It's not going to give you any

22:04 feedback if it crashes. The parser is in

22:07 memory right? So if their KERI instance

22:10 goes down at that point in time when you

22:12 just got the 202 back, you have no idea

22:15 that they didn't parse it at all.

22:16 They've never seen it essentially. So

22:19 for application reliability, we need

22:21 retries on that.

22:23 – When you're talking about “they”, you

22:25 mean the other controllers agent

22:28 maybe not even related to your agent

22:30 other than it's on the same KERIA. So,

22:33 I was experiencing the need for

22:36 retries that I didn't understand because

22:37 I have a good network. But it

22:41 could just so as a defensive programming

22:45 when you get the 202 or other certain

22:47 other areas, you need to retry. It's not

22:49 necessarily a network error.

22:52 Yeah.

22:52 Is that what you're saying? It's the

22:53 KERIA is busy or dropped it.

22:56 Yeah. I mean, so on the 3902 port, the

23:00 CESR endpoint of KERIA, it will just

23:03 take in the CESR stream and put it

23:05 onto the queue of the parser and give

23:08 you a 202 back and then parse it. But if

23:12 it gives you the 202 back and

23:14 immediately crashes, you've now lost

23:16 that data.

23:18 They have no idea that they need to

23:19 parse something. They can't parse it.

23:21 It's lost in memory. But the 3902 port

23:23 isn't used by Signify-ts, right? It's

23:26 I'm not talking about Signify-ts here.

23:28 I'm talking about entire KERI here,

23:30 but agent-to-agent communication. Yeah.

23:35 but also from Signify clients.

23:38 can use it if they want to? 3902 is open.

23:43 Yeah. But Signify-ts doesn't do anything

23:45 with it, out-of-the box, right?

23:47 No, it doesn't out-of-the box. No.

23:50 But also, if your agent accepts,

23:53 an IPEX message, it should

23:56 send it, right?

24:00 Especially for groups. So, for example,

24:03 if you have a group set up

24:07 based on KERIA and Signify

24:10 Let's say, the one person proposes an

24:14 IPX message, sends it to the others, it

24:16 starts collecting signatures. before it

24:19 collects enough signatures, it goes

24:21 down, comes back up, receives the other

24:25 signatures, nothing happens. It doesn't

24:27 send it to the other person. Because the

24:29 granter deck is in memory, the queue.

24:32 So, it's lost it. And the lead of the

24:35 group currently is the one that sends.

24:37 So, a lot of the in-memory decks that we

24:39 use for cues in KERI need to be made

24:42 durable. So

24:45 Is there an issue? If there's not an

24:47 issue with that, I'm going to make that

24:48 one right now. That's

24:49 Yeah, that's bad. That one

24:51 that one's bad, actually. That needs to

24:52 be fixed soon. But you don't see

24:55 these in your normal setups like

24:57 because it's just "happy paths,"

25:00 everything just works. Everyone clicks

25:01 at the same time.

25:02 It's all good.

25:03 Yeah.

25:05 And it'd be nice to have some more

25:07 feedback on the async stuff because

25:08 everything in KERI is async. So,

25:11 if we have better ways of receiving

25:14 results that are more rich on what's

25:17 happening, that would be useful. But

25:20 that's a bit open-ended.

25:22 I mean, if you already have a plan to

25:23 put these into issues, I don't want to

25:24 duplicate it. But is there

25:27 a plan to currently make an

25:29 (GitHub, red) Issue on that because if there's not,

25:30 I'll capture it immediately.

25:31 I have a joint (GitHub, red) Discussion but I don't

25:34 know if it has everything from this. It

25:35 was written a while ago. It's either

25:37 in Signify-ts or KERIA (GitHub repos, red).

25:42 that's some information from Sam on application layer reliability.

25:48 If you want to read it, it's him saying

25:49 that KERIPy is not going to provide

25:52 this essentially.

26:01 What is him saying again? About achieving application layer

26:04 reliability end to end.

26:06 Okay, cool. And the need for retries and

26:08 idempotency, so forth.

26:15 Cool. Finally, well, not finally.

26:25 Especially when you're using group

26:27 multisig, and if you're doing something

26:29 wrong, it can be quite hard to

26:31 understand what you're doing wrong

26:33 because it's a multi-agent setup. All

26:36 your logs are intertwined. There's no

26:38 way to filter on the logs. And even if

26:40 there was a way to filter on the logs

26:41 per agent, there's still too many

26:44 logs, there's no way to filter on the

26:45 action you were trying to do. And

26:49 data again, everything's based on HIO.

26:52 Everything's async, everything's moving

26:54 various ways with little

26:57 observability. It's again fine,

27:00 once you've developed it's fine,

27:01 but sometimes it can be very

27:04 difficult to add new features to KERIA

27:05 and try and figure out where the data is

27:07 actually moving and what's actually

27:09 wrong and why it isn't completing and

27:11 so forth.

27:12 Should the KERIA requests have a

27:14 request ID added to them?

27:15 Probably. Yeah. But it'll need

27:18 some work with KERIPy as well

27:21 because most of the complex logic of

27:24 KERIA is KERIpy. It's a direct

27:26 dependency. It gets pushed onto KERIPy.

27:28 So I need to do some work

27:30 there. So yeah, some instrumentation

27:34 would be useful at that.

27:37 You see you go through the

27:39 logs and you're just trying to figure

27:40 out what's going on, because you've put

27:41 the wrong ID somewhere, and it's ...,

27:44 You can see Kent added some

27:46 stuff here. But we probably need

27:48 to go a bit further so we can actually

27:50 filter on IDs of agents and so forth.

27:53 A structured log.

27:54 Yeah.

27:55 Request ID or client ID.

27:56 Yeah.

27:59 I'll do this one quite quickly.

28:04 Yeah, probably KERI and Signify should

28:05 have some semantic versioning because

28:07 it's an API. And I think you raised

28:10 issues on this. There should be

28:13 compatibility enforcements between them.

28:15 I don't think this is big issue now,

28:17 but in as you move towards

28:20 being a more mature ecosystem and

28:21 getting more developers in, it'll become

28:23 more relevant. Also versioning.

28:28 Right now you have to version on the

28:30 KLI per agent and iterate over them. So

28:34 there needs to be a migration framework

28:36 in place. We've been avoiding this

28:38 for ages because there hasn't been

28:39 migrations, essentially, but when there is

28:42 in production it would be

28:43 problematic. So that's resilience.

28:49 And again these are

28:54 traditional software problems

28:57 Just put some resources in

28:58 and it'll make it a more mature app

29:00 It's not a protocol thing. These are

29:02 all very achievable.

29:05 Finally on Scalability

29:08 And this one is also a kind of

29:11 developer experience

29:14 I mentioned earlier that because a lot

29:16 of the things are async, you need to

29:19 submit like create an identifier, create

29:22 a registry, issue a credential and you

29:23 get an operation back.

29:26 And then you need to pull on that

29:28 status of the operation until it's done

29:30 while it's waiting to collect signatures

29:31 and so forth. That's kind of a pain.

29:35 Typical modern systems are usually

29:38 event driven. Where you would get

29:40 some notification that it has completed

29:42 rather than constantly checking.

29:46 The same goes for notifications. You

29:47 have to constantly check if there's new

29:49 new notifications.

29:52 You know in any general application like

29:54 a wallet, it's going to have to

29:55 constantly check because it doesn't know

29:57 when to expect that you're receiving a

29:59 credential.

30:00 – Single-threaded environment, it could

30:01 completely

30:04 block the rest of the application.

30:05 Yeah.

30:08 We've made event driven in our

30:10 wallet in every application we've built

30:11 because we've just put it in a corner in

30:13 the box and made it event driven from

30:16 the perspective of the rest of our

30:17 application. But it's kind of annoying

30:19 to have to build that every single time.

30:20 But this is very doable, if we were to

30:24 take the query reply

30:26 approach of just using the

30:29 the WSGI response stream, we could open

30:32 up the server ... connection per

30:34 agent. Yep.

30:36 To Signify and do this today.

30:39 I might try that over the weekend.

30:43 Huh?

30:45 No.

30:46 Tell me about it.

30:52 Before I get to that. The reason I put this in scalability is

30:54 because right now you have 10,000

30:56 wallets using one KERI nstance and

30:59 they have 25 pending operations that

31:01 they're pulling, pulling, pulling. These

31:03 are all full HTTP requests. It's way too

31:06 much traffic hitting KERIA.

31:11 It won't work. So moving to something

31:15 like Server Sent Events

31:17 is much better from a developer

31:18 experience, but it's also much better

31:20 from a load experience. And

31:24 it isn't, ..., I've never

31:26 implemented them, so I don't know what

31:27 it's like, but for notifications, it

31:30 wouldn't be too difficult because it's

31:32 just the next notification coming out on

31:34 the event. It's more problematic for

31:36 operations. I explain that in a moment.

31:39 I also Server Sent Events are true

31:42 HTTP, so they're compatible with load

31:44 balancers. So if you start to do

31:46 horizontal scaling and so forth, they're

31:47 compatible, which is why Sam

31:50 would prefer that over

31:51 websockets because that isn't pure HTTP.

31:59 Now the problem for integrating

32:01 with operations is currently

32:02 when you perform an action that requires

32:07 a long run operation to track that

32:09 state of the operation.

32:13 It just stores the related SAIDs and

32:16 metadata that refer to that operation.

32:18 It doesn't have a status of pending or

32:20 failed or anything like that.

32:22 And when you use Signify to check that

32:24 status of pending or failed, it's not

32:26 stored in the database. That's what I

32:27 mean. It doesn't flip in the database

32:29 from pending to complete

32:31 actually if you look at longridden.py

32:35 in KERI, it takes all of those SAIDs

32:38 and compares it against the database of

32:40 the agent to see if it's effectively

32:42 complete. Okay, so it's not checking is

32:46 this operation complete in the database.

32:48 check in based on this operations

32:50 metadata are all the signatures in place

32:53 for this to be considered complete for

32:55 example you can see here for the

32:57 delegations using the 'swain.complete'

33:00 every time you check

33:03 Yeah, basically like a diff on the

33:05 database

33:06 – Yeah, this has been quite a mental

33:08 challenge for me for like two years with

33:10 KERI is the notion that the concept of

33:13 state and state modeling is missing

33:16 and that's intentional

33:18 I think from Sam

33:22 IPEX for example, right, is ..

33:25 I tried to get a state

33:28 machine diagram of all the possible

33:31 state transitions and IPEX of Spurn,

33:33 Grant, Accept, and all that and

33:36 I had a challenge getting to the

33:38 precision but this is similar: You're

33:41 wanting to query the state of an

33:44 operation and it has to be derived and

33:46 has to be requeried to compute the state.

33:48 That is what I think I heard you say.

33:51 Yeah.

33:51 ...of an operation, right? So it'd be

33:54 nice just to get the state back

33:56 from KERIA like 'Oh, it's in Pending'.

33:58 Yeah. So basically while currently

34:00 we can't do event driven with

34:01 Server Sent Events.

34:03 To date?

34:04 Because it's you're checking against the

34:06 database state. Okay. So instead of

34:09 checking the database state or the diff

34:13 when this completes for the first time

34:15 inside KERIPy, it needs to emit an

34:17 event that gets sent on server the

34:19 Server Sent Event

34:21 thread.

34:22 Okay. Because I'm thinking what we could

34:23 do for each operation type is that there

34:27 could be there's already a

34:28 server event connection from the agent

34:30 to the Signify client. It would

34:32 essentially just be a channel to

34:33 communicate. I'm thinking on completion

34:35 it would just send a completion.

34:36 Yeah. But there is no completion of an

34:39 operation in KERIA. It's a state check.

34:42 It tells you it's effectively complete.

34:46 There's no point where it's in the

34:47 database ..,

34:48 – Iit's inferred

34:49 Yeah, it's inferred. So if you

34:51 never ever connect it again,

34:53 nothing changes on the database to say

34:55 it's complete. It's just you coming back

34:57 later and saying this is the status.

35:00 Okay. So because there's no point at

35:02 which that happens

35:04 This is a problem all the

35:06 time, right?

35:06 So essentially

35:08 we need to store the completed state, not

35:10 just derive it.

35:11 Yeah, All of this happens

35:13 in KERIPy so KERIPy needs to tell us

35:16 because this is a KERIA thing

35:18 operations aren't in KERIPy

35:20 – .. locally in KERIA per agent so because

35:22 once we derive it once even though we're

35:24 reading it

35:24 yeah but what triggers to derive it

35:26 other than you're going to move the

35:28 polling inside internally

35:29 – What triggers what?

35:31 What triggers 'Check and it's complete.'

35:34 in KERIA.

35:35 So this is where I think what we should

35:37 do when we have an operation

35:39 There should be a Doer for each

35:42 operation

35:43 once it completes at the end before that

35:46 Doer finishes it should write the

35:47 complete state.

35:48 Yeah. But now you've just

35:50 moved all of the polling internal to KERIA

35:53 and you'd have so many doers for 10,000

35:57 Signified clients. So it would

36:00 be a lot better if internally in KERIPy

36:03 It could just let us know when it's

36:05 effectively completed

36:08 because there's already something inside

36:09 KERIPy tracking this I would say. So

36:11 then it can admit that event or whatever

36:14 and we can capture it and update our

36:17 operation in the database and send the

36:19 event back and Server Sent Events.

36:21 Yeah, I think we're slightly talking

36:22 past each other but I think I get what

36:23 you're saying. I think fundamentally

36:25 when an important event happens in KERIPy,

36:27 it should be admitted saying "this and this

36:29 happened."

36:30 Yeah.

36:32 Rather than having new Doers that are

36:33 also tracking it again.

36:37 Cool, just one more slide. So it's

36:38 actually good.

36:39 No, actually one; so not too bad.

36:44 Yeah, and horizontal scaling.

36:46 People have asked about this before

36:49 I would put this on the lowest priority.

36:50 because other things just need to be

36:52 fixed first to get adoption. But for

36:54 SEDI it would be important.

36:57 Okay, right now you could do something

36:59 like sticky sessions. It would be a very

37:00 naive approach. Where you

37:04 deploy 10 different KERIAs and you

37:06 have a load balancer and once a

37:10 client connects and gets routed to a

37:12 KERIA instance they always have to go

37:13 there, forever. That's like a sticky

37:15 session. It would work now, but it

37:17 wouldn't be great.

37:20 To do true horizontal scale, I think KERI is

37:23 actually already quite suited to it.

37:26 Because all the databases of agents are

37:27 fully isolated and no agents are shared.

37:31 They're just a bunch of LMDB files.

37:34 So if you had to ... [reformulates, red.]

37:40 The easiest thing you could do is have a

37:42 shared volume in your cloud deployment

37:43 that all of the instances can access and

37:45 they would just go to the specific LMDB

37:47 files that they're currently

37:49 serving. And if it's done correctly

37:52 they won't try to access the same LMDB

37:53 file of someone else and you could have

37:55 the agency coordinate that. If you

37:58 wanted more isolation, I guess you could

38:01 separate the volumes, but I think that

38:02 would be

38:05 a little bit complex, not necessary.

38:11 You'd be moving data across

38:13 volumes. I don't like that for now.

38:14 At least I don't think it's necessary.

38:16 Unless there's a

38:20 legal constraint on data or something.

38:24 And yeah, the agency can become the

38:26 load balancer, like it currently is

38:28 basically for the message router in

38:29 KERIA. But instead it would

38:31 actually be outside

38:33 KERIA and route into a more

38:38 light-weight KERIA. So you kind of

38:40 decouple it a bit; be different services.

38:45 I think that's more down the line.

38:46 So that's why developer experience

38:48 and operational resilience are highest

38:50 priority, I think. Yeah, Done!

38:59 – What's this up there?

39:00 That's the file system.

39:03 The agency database is an LMDB

39:05 database. You can see that there was two

39:08 new agents booted. There are the config

39:12 files that get copied in the JSON config

39:15 files with the OOBIs.

39:17 And for each

39:20 agent they have a subdirectory in all

39:23 of these directories. Like there's the

39:25 key store, for example.

39:29 Probably it would be better if you

39:31 restructure it, so that the agent is at

39:34 the top level and it has its databases

39:36 internally so you can more easily take

39:38 it. But if it was a shared volume

39:40 across many instances it wouldn't

39:43 matter.