KRAM IT! with Ari - Ari Argoud

KERICONF26 Day 1 · 37:20

0:00 Ari Argoud | KRAM IT! with Ari | KERI Conference 2026

0:02 Hi everybody. This is KRAM It! with Ari.

0:06 Important exclamation mark!

0:08 placement there.

0:10 So…

0:12 KRAM is KERI Request

0:15 Authentication Mechanism

0:17 and this is for all non-KEL

0:22 events or non-key event messages

0:25 that occur in the KERI protocol and

0:27 some examples include encrypt sender

0:30 sign receiver which if you were at

0:31 Phil's [talk] you'd know is a good way of doing

0:34 Secure APIs to a given identifier

0:38 essentially a replacement for

0:39 shared-secret-based account management.

0:40 You can do exchange

0:42 transactions for negotiating

0:44 contracts or other negotiations. You can

0:47 do credential presentations. And

0:50 so the fundamental problem that KRAM

0:53 solves is ensuring that a message is

0:56 timely and authentic.

0:59 So without KRAM, an attacker can

1:02 intercept a valid signed message and

1:04 then replay that message at a later time

1:07 and you could have a ..., or the

1:10 receiver of that message would have no

1:11 way of distinguishing it from a fresh

1:13 one which could cause all kinds of

1:15 problems, like if you

1:17 were signing on to an account or

1:18 something, you could have someone who is

1:19 able to sign on for you. So

1:23 KRAM works by enforcing a timeliness

1:25 cache window.

1:27 Each message includes a datetime stamp

1:29 and the receiver is going to check

1:30 whether that time stamp falls within an

1:32 acceptable [time] window relative to its own

1:35 clock. If the message has been seen

1:37 before, if it falls outside the window,

1:39 it's just going to get dropped.

1:41 So KRAM exists at the top of the message

1:44 processing pipeline. All of the

1:48 non-key event message flow through a

1:50 non-key event messages flow through a

1:52 single method and KRAM is applied

1:55 immediately after some basic 'allow deny'

1:57 logic if in case you want to filter for

1:59 given identifiers, in case you want

2:01 to not let specific people in. And I

2:04 should note here that KRAM is designed

2:06 for KERI v2 which is not out yet. [update July26: v2.0 is out]

2:08 But it represents a change from a prior

2:10 approach wherein before we would escrow

2:13 messages when authentication

2:15 information wasn't immediately available

2:17 like if we didn't have the Key Event Log

2:18 whereas now we just drop them and notify

2:21 such that the Key Event Log can be

2:23 retrieved via something that is out

2:25 of scope for this talk.

2:29 So before diving into KRAM's

2:31 mechanics I just want to make sure that

2:34 everybody is familiar or largely

2:36 familiar with these

2:40 few key terminology points. Does anyone

2:42 have any questions about KERI? Well,

2:45 okay. Questions about KERI might be a

2:47 little broad, but we all know what AIDs

2:49 are,

2:51 KELs,

2:53 Key State, we've talked a lot about it

2:54 today. SAID is maybe worth touching on,

2:57 which is the Self-Addressing

2:59 Identifiers, which are digests computed

3:01 over the contents of a data structure.

3:03 And basically, this is an ID for a given

3:06 message or in this instance, it is or

3:09 referencing it as an ID for a given

3:10 message. – So it's a context

3:13 identifier based on your payload?

3:15 – Yeah, that's that seems

3:18 correct.

3:18 – It's actually the whole message,

3:20 right? – It's a digest of the message.

3:21 – It's a digest of the message?

3:23 – But it's embedded in the message…

3:26 […] protocol. So that there's one universal

3:29 identifier for that message.

3:33 – So it's unique per message even if

3:36 applied in a thread.

3:38 – Yeah.

3:39 – Yes. – Could you repeat the

3:41 conclusion of this because you are the

3:44 only one with a microphone.

3:47 – Oh yes. So basically

3:52 We have SAIDs as unique identifiers

3:54 which are digests of a given message

3:57 such that no two messages that are the

4:00 same would have different sets. I.e. if

4:04 two messages have the same set

4:06 then they must have identical contents

4:08 and time stamps. So then Seals,

4:12 briefly is a cryptographic commitment

4:14 which can anchor arbitrary data to a

4:16 tree of hashes or to a particular

4:20 event in the key-event sequence. And

4:23 we're assuming with KRAM that the

4:26 receiver will already hold the copy of

4:28 the sender's KEL. Which is again

4:30 kind of gesturing towards that fact that

4:32 that retrieval mechanism is out of scope

4:34 but it does exist or will exist.

4:38 So in KRAM we have three authentication

4:41 types. We have an Anchoring Seal

4:44 Reference which is a message that's

4:48 authenticated via a seal in the

4:50 sender's KEL. And this works because

4:53 you can preemptively authenticate your

4:56 messages as a sender by digesting

5:00 them into a seal with the

5:02 authentications already

5:05 created and then adding that to your

5:06 KEL, and then continuing

5:10 your Key Event Log. I think this would

5:11 be done via an interaction event and

5:14 then other people can just look up that

5:16 seal and confirm that is a

5:17 pre-authenticated message. We have

5:20 Single-Key Signature which is just one

5:23 key in the key list. I should mention

5:27 that, let's see…

5:31 when we are determining this type of

5:34 authentication we're looking at key list

5:37 cardinality.

5:39 So sometimes an identifier will have

5:42 multiple signing keys but it won't be a

5:45 multi key signature and for the purposes

5:47 in .., eventually that matters, it is

5:50 covered in KRAM because KRAM does

5:51 authenticate but what we're doing here,

5:54 why we're breaking these

5:55 things out is because…

6:00 … we're trying to

6:02 do Amplification and DoS (Denial of Service) mitigation.

6:05 So, it's important to do the cheapest

6:09 method of dropping messages before we

6:13 get into the more expensive

6:14 authentication-based

6:16 methods of dropping messages.

6:18 So, first we would…

6:23 first we're going to look at the key

6:24 list cardinality. We're going to say is it

6:26 a single signature? Is it a multi key

6:28 signature? Does it have a seal

6:30 reference? And then if it's multi key

6:32 and it has a seal reference, we're going

6:34 to try and look at that seal reference

6:35 first before we go on to see if that

6:38 multikey has some has some validity

6:41 to it, if it has signatures that are

6:43 valid and validating those.

6:48 All right. So we'll talk briefly about

6:51 multisig. You guys have probably heard

6:52 it all already today. But this is

6:56 when multiple signing keys are required

6:58 to authorize an action. These would be

7:00 distributed signing keys. In

7:03 practice, that's when a multisig group

7:05 sends a message and not all the

7:07 signatures may be attached at once. The

7:09 different members of the group may sign

7:10 at different times and from different

7:12 devices. The receiver needs to

7:15 collect those signatures,

7:18 needs to collect those signatures,

7:21 sorry

7:24 incrementally until we meet a threshold

7:28 which will validate that message

7:34 if the message is able to collect all

7:37 those signatures.

7:39 So this is where we have different

7:41 KRAM [time] windows. KRAM uses a long lag

7:44 window for multi key signatures wherein

7:47 as opposed to a short lag window for

7:49 single keys and seal reference

7:52 messages. The long lag might be minutes,

7:55 might be hours or days whatever is

7:58 appropriate for the operational tempo of

7:59 the multisig group and the short lag

8:02 only accounts for network latency and

8:04 initial processing times. This is that

8:06 window that I was mentioning earlier.

8:08 It's going to be different depending on

8:11 the authentication mechanism that we

8:12 suss out in the beginning of the

8:15 KRAM-It method, which is why this [talk] is

8:17 called 'KRAM It with Ari.'

8:21 So,

8:27 let's talk about what a Replay Attack is

8:30 I mentioned this already but I wanted

8:32 to go through it again just so that we

8:35 can

8:36 create some context for the next

8:38 things that we're going to be talking

8:39 about. So, Replay Attacks are

8:42 fundamental threats in any authenticated

8:44 messaging system. The concept is that

8:47 an attacker will intercept a message

8:49 that was legitimately created and signed

8:51 and then is able to resend that

8:54 replaying it to the receiver at a

8:56 later time. If the receiver has no

8:59 mechanism to detect that the message is

9:02 a duplicate, it may process it again as

9:04 if it were fresh. And depending on

9:06 the message type, this could grant

9:07 unauthorized access, trigger

9:09 duplicate operations, or otherwise

9:12 cause harm. So, it's important to note

9:15 that the attacker doesn't need to forge

9:17 or modify anything in the case of a

9:19 replay attack. The message is

9:21 genuinely signed by the real sender and

9:23 the attacker can exploit the

9:26 fact that message has a valid

9:27 signature on it. But that doesn't

9:29 mean that the message is provably

9:32 delivered right now.

9:34 So this is why timeliness is a

9:38 core mechanic for KRAM. A message could

9:40 be properly signed but it must also

9:44 fall within the timeliness window that

9:46 is able to correctly filter out

9:52 any messages that may

9:54 be sent by malicious bad actors.

9:59 – In current-use protocols, how is that

10:01 mitigated?

10:03 - OAuth2 or other examples

10:06 – I don't know. I've only studied

10:09 KERI, so I couldn't tell you how OAuth 2.0

10:12 does it.

10:13 [inaudible] approach in most secure channel

10:16 – is it a nonce

10:17 – by TLS is you do a challenge response.

10:22 So it's you […] with three messages. So you send

10:25 a request, the host sends you back a

10:30 nonce, you sign the nonce or encrypt the

10:33 nonce, send it back and then they decrypt

10:35 it or verify the signature and that

10:38 nonce is basically

10:41 not replayable as long as they don't

10:43 ever generate the same nonce and accept

10:45 it. That means they either

10:49 assume an ephemeral secure session so

10:52 they don't have to save the nonces but if

10:54 you want something like multisig

10:58 that can span multiple sessions then

11:01 your host because now you have to store

11:02 all those nonces forever get replayed.

11:06 So, when you look at Replay

11:08 Attack Protection, when you go away from

11:11 ephemeral interactions to things that

11:14 have to last for long periods of time,

11:16 most of the existing mechanisms just

11:18 don't work. They just

11:19 don't scale. You have to remember things

11:22 forever or you have to start to build

11:24 provable caches.

11:26 – Statistically random with large

11:29 – Yeah,

11:32 – That makes sense actually.

11:34 – Okay. – I want to ask

11:35 to summarize.

11:35 To summarize… other authentication

11:40 mechanisms like OAuth 2.0 do multi-

11:44 message authentication via nonces which

11:47 are not scalable as such.

11:58 Now,

12:00 talk about some fundamental protections

12:02 of KRAM.

12:04 So, the first thing that we're going to

12:06 do when well, we're ordering

12:08 these by the cheapest means of

12:12 making sure that your message is

12:14 authenticated per the given point

12:18 in the protection that you're at. So the

12:20 first thing that we want to do is we

12:21 want to look up this

12:23 message. Like we said, it's got a SAID on

12:25 it which is unique. So we're looking it

12:27 up against a cache. And if that message

12:31 is there in that cache already, then we

12:34 know it's a duplicate because no message

12:36 can have the same set as a message

12:38 that's already existed. So you know

12:39 you've got something funny going on if

12:41 you're… or well it could be

12:44 it could not be something funny. It

12:46 could also just be someone who's

12:47 legitimately trying to send a message

12:49 again. But the point being that the

12:51 first thing we do cache lookup cache

12:53 lookup and drop if we already

12:57 have something there. We are also

13:00 checking to see if these are multikey

13:03 messages because we can have a message

13:05 that has the same SAID but has

13:09 additional signatures that must be

13:11 collected towards a threshold. So there

13:13 is a little bit of nuance with respects

13:15 to multi-key

13:19 messages.

13:21 Next, you're going to do a timeliness

13:23 check which only occurs for uncashed

13:25 messages. And you're going to do this

13:29 nice little formula here. This

13:30 is the real magic behind KRAM or at

13:32 least some of it which is a receiver

13:34 date time minus drift minus ML which is

13:40 going to be a…

13:44 this is a variable, this could be

13:47 different depending on whether you have

13:49 a multi key, it might be a longer lag or

13:52 whether you have a zip wherein you're

13:56 going to be waiting for an even longer

13:58 time to do different negotiations of

14:00 contracts or multiple messages within

14:02 within a given very long time window.

14:06 So right, do you see here ML equals SL, i.e.

14:10 we're doing for seal reference and

14:12 single key we swap this out for a short

14:14 lag, long lag for multi key, etc.

14:18 Next you're going to try to

14:19 authenticate. you're going to resolve

14:21 the auth seal type by checking the seal

14:24 reference first because it's cheap and

14:25 then you're going to go on to signatures

14:28 and with single key you just have to

14:31 validate cache and accept. Multi key,

14:35 again, you're going to verify the

14:37 available signatures and then you're

14:39 going to take a look at that threshold

14:41 And see if it's met at that point.

14:43 And if not, then you're going to

14:47 you're going to well, you're basically

14:49 just collecting those and waiting for

14:50 more messages to come in until that

14:52 threshold is met. It's important to

14:56 note that the prune [time] windows always

14:59 must be greater than the accept windows

15:02 which is going to prevent gap attacks

15:04 during configuration changes. And again

15:06 we're doing… Yes?

15:09 – The window is a measure of timing.

15:11 – Yeah, it's a sliding

15:14 window. So a given window will be an

15:16 offset in one direction and another.

15:18 Right. So like if it's what time is it?

15:21 Three. – Or wall-clock time?

15:24 – Say that last part again.

15:25 – Is the expectation that it's wall-

15:27 clock time or something else like tick

15:29 counts?

15:30 – It's going to be wall-clock time,

15:31 right? Network time.

15:33 – Network time.

15:34 – Yes. So…

15:37 – So is there a dependency to have secure

15:39 time?

15:41 – Yes. Well, no. Okay.

15:43 – No.

15:44 No. No.

15:45 – What's secure? What do you mean

15:46 secure time?

15:46 – Time is always relative to the host. So it

15:50 doesn't matter what an attacker

15:52 does. If they're not

15:54 synchronized with the host time, they

15:57 can't lie about the time, right? And

15:59 and if the host records the timing, they

16:02 can't do a clock replicate attack on

16:05 it. So that when they spin back up, you

16:07 know, they can't like push the time

16:10 back, you know. – So time, but the

16:14 security is the

16:15 the host itself. And the

16:17 assumption is the attack [inaudible]

16:23 – So to be clear…

16:24 – I don't know what you mean by secure clock

16:27 but those [inaudible]

16:30 [inaudible]

16:33 – Yeah. But in this case server you just

16:36 you are the server.

16:37 – Yeah. You are the server.

16:38 – Okay. So you are the server which is the

16:41 security essentially to answer your

16:43 question. Okay. That for the microphone.

16:47 Let's see. Did I miss anything on

16:49 here?

16:51 Oh, right. Key State Protection.

16:54 It's good to be aware that this is for

16:57 a single given key state. And if

17:00 sender key state changes in the duration

17:03 of a given, you know, messaging

17:06 exchange, like if you're collecting

17:08 signatures on a multisig and the sender

17:10 key state changes at some point during

17:11 that, then that messaging operation

17:14 is going to have to happen again.

17:17 – You said gonna have to what?

17:19 – It's gonna have to happen again.

17:21 – Oh, yes.

17:21 – Yes. Okay, good. You're looking at me. I was

17:24 like…

17:25 – I didn't hear what you said. I didn't

17:28 hear the…

17:29 – Thank you. Okay. Now we get to talk

17:32 about the cool stuff. So, it's all cool

17:35 stuff. I'm just kidding. These are

17:39 kind of… they're not introduced by KRAM,

17:42 but having KRAM as a Replay Attack

17:47 Defense introduces other modes of

17:51 exploitation. So we have gap replay

17:53 attack and gap first play attack.

17:57 Essentially…

17:58 I'm actually going to move on to the

18:00 next slide because I have some nice

18:01 visuals here for you. For a gap replay

18:04 attack. Here we have our timeliness

18:07 window. We have our message being sent.

18:09 It's accepted and cached and then the

18:11 message gets pruned. You have an old

18:14 timeliness window right here. And after

18:18 the message is sent, before it's

18:20 accepted, a bad actor is going to go and

18:22 grab that message. Right?

18:25 Next, the timeliness window expires. The

18:28 message is pruned, but then you as the

18:31 server or the receiver is going to

18:33 change your timeliness window and

18:35 increase it. So here after the message

18:38 is pruned a bad actor can replay that

18:41 message in this little vulnerable gap

18:44 right here. We're going to talk about

18:45 how we solve these with KRAM but these

18:48 are complications that are introduced

18:50 that need to be addressed.

18:52 – You dynamically change [inaudible]

18:53 – If you dynamically change windows

18:57 – Yes exactly. So this is when you have a

18:59 running system and you want to change

19:01 your timeliness window for a given

19:03 message type or other configuration.

19:07 So for a gap…

19:11 hold on… sorry this one should say gap

19:14 first play visual. This is a gap first

19:16 play attack. Basically what happens

19:18 here is you have a message sent. A bad

19:21 actor is going to intercept that

19:22 message, but it's going to come through

19:24 with the receiver of the message

19:26 never having touched it or looked at it

19:29 because it comes through after your

19:31 timeliness window.

19:34 Post that. This is a little bit

19:35 of a confusing time scale, but post that

19:38 you're going to increase your timeliness

19:40 window. And then that bad actor who's

19:43 saved that message can first play it

19:46 into your new increased timeliness

19:49 window. – What's the intuition for why

19:51 they increase third time in this window?

19:54 – Could be any number of reasons like

19:56 let's say you're doing a multisig

19:59 exchange or a multisig and the

20:02 group participants are

20:04 not getting their signatures in

20:08 on time for given messages and you say I

20:10 want to give these people a little bit

20:11 more time. You increase the window but

20:13 you also open yourself up to this…

20:14 – Just on a network with lots of packet

20:16 loss.

20:17 – Exactly. – The network starts to have

20:18 congestion and now your latency [inaudible]

20:22 – So [inaudible] could simulate that using the

20:26 Denial of Service attack somewhere in the network

20:28 – and then [inaudible] it into the

20:31 window.

20:32 – Yes, so to reiterate

20:34 another reason that you could

20:36 encounter this or that you could

20:38 encounter needing to change your

20:39 timeliness window is if you have a lot

20:41 of network latency. Perhaps an

20:43 attacker is DDoSsing you or maybe you're

20:45 just scaling up and you realize that you

20:47 need more time to process. So

20:49 anyways, now we're going to talk about

20:51 how we address these. So

20:56 … oh… now this is the right slide. We

21:00 have three cases.

21:02 If you're decreasing a [time] window size,

21:04 that's safe and you can apply that

21:06 immediately. If you're increasing a

21:08 window size, you need a two-step

21:10 process. And if you're changing

21:14 cache type granularity, I'll talk a

21:16 little bit about configuring cache types

21:18 in the later slides. That's more complex

21:20 and would require a worst case analysis,

21:22 but that's a little bit out of scope for

21:24 this. Anyways, the two-step

21:27 processes for increasing it is you're

21:29 essentially going to want to have a

21:32 prune lag wherein

21:35 when you do an increase, you have an

21:38 extra lag where you're not pruning

21:40 windows until you've exceeded the time

21:43 delta between your old timeliness window

21:45 and the new increased timeliness window.

21:47 And that actually covers both gap replay

21:49 attacks and gap first play attacks.

21:58 Then we have a little bit about the

22:00 configuration

22:02 wherein this is designed for

22:05 incremental adoption. So not every

22:08 deployment is going to be able to run

22:09 KRAM for all messages types overnight

22:12 and the configuration system

22:14 reflects this. So, we have up here the

22:16 'enabled true' which just means your KRAM

22:19 is on. Then you have denials and

22:22 these are each keyed to KERI

22:26 version. So we have version one and

22:30 version two right here. And these are

22:33 going to be message types. And these are

22:36 going to be…

22:40 These are going to be…

22:49 route. Yeah, sorry.

22:53 And the denial systems supports wild

22:56 carding through empty fields. So these

22:59 one comma zero KERI version one wild

23:01 card wild card just means we're going to

23:03 deny all KERI v1 messages for KRAM

23:07 or we're not going to use KRAM on

23:12 KERI v1 messages, right, and for these

23:15 we're not going to use KRAM on

23:17 KERI v2 messages with a pro message

23:20 type and then any route and for the

23:23 third one we're not going to

23:27 feed messages through KRAM for KERI

23:30 version two reply messages with a

23:33 “/end” route.

23:36 And the denial logic is applied at the

23:39 top, before we even get into any of the

23:42 message processing stuff that we

23:45 discussed earlier.

23:48 That is pretty much all I have. Well, I

23:51 was also gonna say

23:54 that

23:58 a lot of people worked on this. Sam

23:59 wrote the white paper. Several other

24:01 people from the KERI Foundation have

24:02 put a lot of effort into getting the

24:04 implementation to where it's at now. And

24:06 I know that many other people are going

24:07 to work on it before it's complete. So,

24:10 I give them a round of applause.

24:15 Does anyone have any other questions?

24:18 I run a little short.

24:23 – Yes.

24:23 – Where do you see this being put into

24:26 application within, I mean, let's just say

24:29 the healthKERI

24:31 platform.

24:32 – So this is going to be like for say

24:37 for example for our account

24:39 communications with ESSR they could use

24:41 this. For any credential

24:45 presentations they could use this, right.

24:47 So, we were talking a little or we

24:50 haven't talked about this yet, but

24:53 actually… I'm going to abstain because

24:54 I'm not sure if I'm supposed to say that

24:55 yet, but those two use cases are

24:58 pretty relevant to us as healthKERI.

25:00 – Absolutely.

25:09 – Yeah. [inaudible]

25:12 – Yes.

25:13 [inaudible]

25:17 – Yes. So, IPEX EXN messages, all covered

25:20 by KRAM.

25:23 – Yes.

25:23 – I just want to understand the problem a

25:25 little bit better of why we need to keep

25:26 this just I think that's what he had

25:29 asked.

25:29 – Oh, sure.

25:30 – I understand it replaces the nonce, but

25:32 why is that beneficial? Like what?

25:36 – What?

25:37 – Yeah. Just so if you have a nonce I

25:40 assume it's kind of like a temporary

25:41 token.

25:42 – Yes.

25:42 – That's what I've used in the past. That

25:44 then it would increase the database size

25:47 or maybe you have to have

25:50 connectivity back to… I'm trying to

25:51 figure out why.

25:54 – It's like an elegant way of solving

25:57 replay attacks without attaching nonces

26:01 to every message because the time itself

26:03 functions as the nonce,

26:06 like the time attached to the message.

26:08 It's not that we don't

26:10 have a nonce the time is the nonce right?

26:12 – [inaudible]

26:13 interactive authentification right

26:15 [inaudible]

26:18 so if you think about

26:20 network latency we're doing at scale

26:25 requests

26:27 every request

26:30 has to create a TCP connection because

26:32 it has to be a secure channel

26:35 otherwise you can't control it, so

26:37 all asynchronous messaging is off

26:39 the table. This will work

26:41 just as well with asynchronous

26:43 messaging as synchronous messaging.

26:45 It'll work just fine where the

26:47 sender and receiver use a completely

26:49 different channel. Doesn't matter.

26:51 Because there's no interaction, right?

26:53 So now you can get rid of all of those

26:54 constraints and then the latency for a

26:58 request is zero because the request

27:01 either goes through or it doesn't.

27:04 Whereas if it's interactive, you send a

27:06 request, you get a reply, you send a

27:08 response, and now you've

27:11 got three passes across your network,

27:13 which if your network's loaded or

27:15 congested, you triple your latency

27:18 always. And when you go to scale, that

27:20 doesn't scale well. And you know,

27:24 so that's a reason not to use

27:27 interactive.

27:29 Two reasons, one asynchronous messaging

27:32 …and remove

27:35 the latency at scale.

27:37 – So to reiterate for the microphone:

27:40 this is asynchronous which traditional

27:44 nonce-based authentication is not and

27:48 this also works at scale without having

27:52 to rely on the traditional mode of

27:54 exchanging several messages in order to

27:57 meet the same

28:00 functionality.

28:02 – So [inaudible]

28:04 statement, right? You know what's what

28:06 was the difference between KERI not

28:08 having KRAM and now KRAM existing and I'm

28:12 just maybe restating but

28:15 […] restated for the replay attack but so

28:18 was that a vulnerability and existing

28:20 vulnerability right now?

28:22 – Yes.

28:22 – KRAM has been proposed for quite

28:25 quite some time but nobody thought it

28:28 was important to…

28:32 – So, in our authentication server, we

28:35 have a custom build replay [inaudible]

28:38 and so we can now get rid of that once

28:40 we upgrade KERI 2.0.

28:42 – Yeah. So, that was traditionally not

28:44 covered by the KERISuite or KERIPy…

28:46 [inaudible]

28:56 We want it to be something

28:58 that everybody gets for free. – Awesome.

29:09 [applause]