Stop reacting. Start resolving. Introducing OrionIQ, bringing agentic observability to your stack. →
Webinar: Agent-to-Agent Communication: The Future of Automated Incident Response – June 9 @ 12:00 PM ET / 12:00 PM IDT
Hi, everyone. Welcome to the our next part of our AI observability podcast series. Today, we'll be talking about the process of automating the NOC at Thetaray Maybe let's introduce ourselves. I'll start. My name is David Lotan. I'm Log's VP product. Galy? Hello, everyone. I'm Gali. I'm a enterprise account manager at Logz.io. Hi. My name is Alex Lushenkov. I am SVP of customer engineering at Thetaray. Hi. I'm Yossi Cohen. I'm the senior and I and ops engineer at Thetaray. And good morning. Good morning. Good morning. Thank you. Pleasure to have you all here. Thank you. It's our pleasure as well. Great. So we've been a customer, we've been our partner for several years. What a journey. So do you want to share with us more about what Thetaray does? So Thetaray basically, we are fighting financial crimes. We have a unique model, AI model, that knows how to detect financial crimes for financial institutions. Are backing up the largest institution around the world in Europe, in US, in LatAm, in Africa, and other parts. And we are very proud of our solution because we're also doing good for the world, not only for ourselves. Good. Guys, everyone is talking about AI, but a lot of organizations are still stuck at the chat phase. What was your experience using AI agents before before Logz.io? So I think, just to give you a context, a year ago, we initiate we initiate a cross organizational project called Nova. Nova basically was how we can embrace AI and use AI in order to improve our operational efficiency For onboarding of our customers, for other workflows inside the company as well. And once we started, it was contagious because everybody wanted also to implement more and more AI solutions. And this is how we get to their support workflows, and this is how we started these initiatives. Because before what we're going to discuss today, we're going to we had auto NOC, okay, with eight students, I think you'll see correctly. Yes. 20 fourseven. And 20 fourseven. Basically, the model was they are looking at the screens, they're trying to find anomalies and try to find the issues. And, you know, once a human eye looking at the screen, it's room for mistakes. And I think it was very obvious from the beginning, this is the place we want to be. And this is the first thing that we want to automate inside the company on the support level. Interesting. Yeah. Maybe you want also to share more about what's your day to day team activities? So if we're talking before the implementation within the NOC and within the support team, we have a concept in our own product that is unknown, unknown because we are an unsupervised machine learning platform. And it was a hard work first to react to a lot of the monitoring tools and scales and tools that we've implemented, Logz being the center of it. And it was really hard work to try and figure out where are the failures, where are the errors are coming from, what did we not think about, and that kind of activity. When we launched the NOC with the students, we got better at it because we had more eyes, more observability, more ability to reproduce issues. We have a very complex platform, so trying to grab everything all at once is a little bit challenging. And as we progressed into the idea of Nova and starting to implement it, we started a little bit tipping our toes in the waters and checking temperature. And we saw the benefits or we saw the potential of what it would be if there would be an automated process that does the reasoning for us, okay, and knows how to go and find those unknown unknowns. Again, how long was it that the unknown? So we have started testing different approaches about a year ago. A year ago, just started. When Nava started and launched, we were a little bit the run to the litter because of a huge backlog and because of the need to constantly retrain. We had eight people in the team. We had a core four people that stay stuck with us, and then there was the other half that, by nature, just were replaced. Those were students in The tuition rate was very, very high. Yeah. Because students typical Yeah. Because they're coming, they want to work for, what, two, three months. And they leave. And all sorts of different reason with the reality that we had in the last three years. So people are calm, people go, people go to reserve. You need to preserve knowledge, and that was a little bit challenging. So we approached the AI with it. We were very conservative and very worried about it because we are responsible for so many transactions that pass through our systems. Those are millions and millions of dollars and euros and whatnot. But we saw potential, so we decided to plunge in. The water was not that cold to start to test it. Quite the opposite. Yeah. And this was pushed by your leadership, right? Yeah. So who came with the idea? So basically, a year ago, we started on the CEO level and the board level to push AI across all the companies. This is the project I told you called Nova. And part of that was first of all, we decided to concentrate on the onboarding of our customers, okay, and to see how we can improve and be better there. But immediately, it was like it was contagious. Everybody in the company wanted to do AI and to so we have today our finance people, our HR people creating amazing tools for ourselves, and it's really great to see. And it was natural to us that the next place to look and where can we improve ourselves is the NOK, okay, and our support. And this is how we got to this point. I want to say that we are working with many customers, and I think that we were so impressed on the Logz side of the seriousness ThetaRay had to this project because tons of companies obviously are trying to focus today on AI. A lot of them talk. You were so serious. And again, we came back and said, guys, these are the perfect partner for us. So this is from our perspective. I couldn't agree more. I'm working with team for you've been a customer for us for several years, and I know by the nature of your business that Thetaray is very mission critical for your customers. I remember we also spoke about how important is the production, stability, and the performance. Do you want to more share from your perspective how you Yes. Saw So from our point of view, we are, as I mentioned, we are backing up most important financial institutes around the globe. And for us, we cannot afford ourselves to be down or to create the issues to our customers because basically, it's money. Okay? If we are down, so somewhere around the world, somebody will not get his money into his account. Okay? So for us, it's a critical mission critical product. I think the idea for our partnership was, at least from my perspective, okay, and this is how I saw it, I wanted to move from reactive mode to proactive mode. I think this is the greatest achievement that now we are seeing those results in production because, Josef probably can share more information, but basically we are getting much more alerts. Okay? We have much more visibility for our production environments. Yes, we're creating more tickets. Okay? But I think basically it's the tickets that we missed while using the students mode. Okay? Because eventually, you know, when human looking at in in middle of the night and, you know, here it's agent that never sleeps. Okay? Always awake. Okay? Always know how to respond and what to do. Okay? It's still we're not we are still in fine tuning of this product, but I think basically we already see great results, okay. And we want to shift our manpower to be proactive. Instead of just looking at the alerts and trying to figure out what happened there, our goal is that this information will already be chewed to our support engineers. Okay? It will be provided to them exactly what happened, and they can be much more proactive. I think last week, we have an issue in our US region in auto NOC. Find it in the second. We notify the customers. We show proactiveness. And I think Yeah. Those those are great results. So the chase was always beat the customer to the to the event. Okay? You don't want them complaining. You want them being Absolutely. Okay. And then if I go back a few minutes into that conversation, when Nova started, our mindset was so conservative, so weary of production. We were so protective of it, okay, that nobody else is allowed to touch it. Nobody else is allowed access. We are the the guardians of it, and we are responsible for it by decree within ThetaRay. Okay? This is was my I I had the keys to the city for a long time. So the And he has a little of a bodyguard. Yes. Don't miss. Do not miss. Okay? You touch my production, we are going to meet somewhere dark. No. We're kidding. So you you have to understand that the mindset that we started to speak about AI about a year ago, no matter what was the solution within support, people thought about, okay. Maybe an agent that will help me write the RCA better. Maybe an agent that will help me deal with the backlog of the tickets just to reorder and prioritize them. Nothing tangible that would touch production. Yeah. And that is important to understand because this is what blogs did for us. It allow us to break that psychological or sociological sociological barrier within our mindset and to think big and bigger, okay? Still conservative because we test everything rigorously. We don't allow anything AI to do still things automatically before a human in the middle comes in and reviews it. But we have gotten better in the detection. We have got better in the communication, which is crucial within our field. You have to understand that in some cases, our mindset is that Thetaray is a post processor platform. So the speed is as much as important as the accuracy and the detail. It's not the top thing. It's equal. And then in other places, it's part of the processing of the transaction. So the compliance and the speed are both very, very, very important because you don't have time to waste. You are in milliseconds. Yeah. Understood. So that communication, that alerting, that ability to detect at least where the geographical area of an issue is, that is crucial to us. And we have expanded. We went live three months ago, and we have expanded the ability just within this phase, I think, in about 40% for that matter. Yeah. So I can like, we've been part of this journey, and I know that you did this very gradually, very, like, very safe. At times, we wanted to, you know, pedal to the metal, but, obviously, we understood completely that you needed to build trust. Absolutely. And by the way, this like, I know that we've been working with Heteroi for a while now. I know that you're very, very KPI driven. How did you measure NOK? That was actually easy because NOK was established, the team itself, KPI based. We had to approve have to understand that the NOC itself as a team was established two years before the implementation Automatic NOC. So we had to prove it from day one. We actually managed to prove it two weeks afterwards. We had the leak and the of water into the chief electric board at the eve of Passover two weeks after it was. So having a NOC in the office Help. Proved itself. Yes. Yes. Yes. Our amazing admin made it happen, like, in the middle of Passover to replace the entire electric bolt. So NOC proved itself, but we measured it. You never told us that. You never told us that. Many stories you never told us. Yeah. No worries. So if I look at it from that from that perspective, we were measured constantly. We were measured about the response time as FERT and about the ability to bring a case eventually full circle into resolution time in a shorter time. We managed to reclassify cases better based on their priorities and based on where they are or what they are because customers tend to complain and not to report, which is a very natural thing. Yeah. Absolutely. We were measured by KPIs on those people from day one. So turning over that KPI into something that we measure for the accuracy, speed, relevance, detection for an automatic NOC That was something that kinda came natural to us. We set the goals. We set the time. And we also wanted to understand that if we start from a certain point, for instance, I think we took a 126 playbooks that we had, playbooks and runbooks, and then we turned over about 60% of them into agents in Logz.io That react to certain types of alerts. We then expanded more, but that was the base or the knowledge base that we had that we turned over. We started from a certain grade and from a certain accuracy that we got. And we measured throughout time how the accuracy became better and better and better. And we measured the time that it took us to get there. Now, if it would have taken slower, like much, much, much slower, we may have had a benefit in the middle. If it become even faster, we may have been a little bit suspicious of it. But I think the pace was both dictated by us of what we were winning or what we were able to feed back into the agent and feed back into the model, but also was natural to us in terms of the flow and the support that we got. So that, alongside with already a substantial foundation of KPIs that we had, was very, very, very helpful for the success of that of that phase. I also think that, yes, we had set of KPIs, and, of course, we use them together with you. But we're in the middle of the process. We're becoming we became more hungry, and we added more KPIs. Okay? And I think because when you see the power of the platform and you see what can it do, I think you become it opens your mind. Okay? If you ask me, everybody now, afraid of AI, it will take our jobs. And I'm saying no. It will create more jobs. It will create more opportunities. Okay? If for now, the barrier was the execution level, okay, how we create the code, it's no longer a barrier. Okay? And now the execution is something alternate. Yes. We still need people in the loop. We need to see the code is validated in grid, but I think now the barrier is our imagination. Okay? Yeah. And once we are there, we can create new set of KPIs. We can be much more aggressive. And as I mentioned, our goal together with you is to move from reactive to proactive. To more. Yeah. Proactive. Proactive and also to active from that matter. I wanna say something because we mentioned the people and we mentioned all of it, and there's a lot of fear that it's gonna take all of that. We took a team of eight people with a team leader that worked part time. On the net scale of it, we didn't lose a single position. Okay, we may have relocated one because we wanted to cover by humans to be more full of the sun. But eventually, if you scale the full time jobs that we announced since then, we took that team. We let go of people who wouldn't go and stay with us for long. And then from the leftovers I don't want to say leftovers. The people who were left, we eventually chose the people that showed the most level of engagement and basically interest in staying long. I think out of the four, one decided to leave, one became a QA, two remained the core team in Israel. So we took eight, and we turned it into a four teams, a team leader and three people, two in Israel, one in Latin America. We didn't lose the net jobs, basically. We just turned them over. So that fear that we have is natural by everyone, but I don't think that it's actually coming for our jobs. I think it is going to reshape how we do things. Yeah. And now we have a better, more professional, more focused, more mission critical oriented and goal oriented team of people who are junior SREs. Understood. Understood. So you mentioned, Alex, you mentioned it previously, and Yossi, also you mentioned, like so it was a strategical decision for the company to move to AI. Still for organizations that do have NOK driven by people yet, what could you mention what are the the biggest pain points of I think having a team there? Yeah. I think the main pain point now, as I see it, is the as you mentioned before, is the barrier of fear. Okay? Because now everybody, you know, what you know, can I trust the agents? Can I okay? Will agents do the same job as humans? I think agents doing much better job than humans. Okay? Today, we are detecting stuff that we never detected before. Okay? And it will allow us to be more proactive and provide a better service to our customers, and it's it's a win win for everybody. Okay? I think also we can shift those this team of four people, as Josu mentioned, now do more stuff instead of just looking on the screen. Okay? And detect stuff. Detection is done by agents. Now you need a human to understand if it's a real issue, what needs to be done, how I am imagining this situation right now. And for that, we need smart and bright people. Okay? And all the, let's say, dirty work is done by agents. Yeah. Yeah. 100%. Can you share more about Alex, what led Thetaray to to working with Logz.io? So I think for us, it was a very natural decision because you already been there as a partner for a long number of years. Okay? You already have the access to the logs. And when we started to look for the solution in the market, so for it, like, for us, it's, you know, it's obvious. Okay? Because eventually, once you have access to the data, okay, it's very easy, okay, to build the agentic solution as you build for us. Okay? And we are design partners on that and we are very proud of this solution. So by the way, we also have our own agentic solution called Ray, okay, that also detects stuff and creates analysts and then inside financial institution to become much better analysts and provide much better support. So for us, once you have the access to the data, okay, it's very easy. I'm not saying easy because you invested too much on that, but I'm saying like Actually, it was mutual. Yeah. So but once you have the access to the data, you can understand and you can build solution much easier from somebody who doesn't have that You understand the logs, you understand your business, you understand our business. Okay? And I think for us it was a win win. Yeah. Thanks. I want to touch on what you mentioned earlier, Yossi. You mentioned that Thetaray is very and specifically, the team is very KPI driven. When you look about when you looked about the right partner, what you were looking for in in the right partner and how it met your expectation in what you are trying to achieve? Attention. Attention to details, understanding. I think that when first of all, David came to us and he begged. Okay. You're saying, no. No. Thought we agreed. We all praise this. No. I'm kidding. But David and Miles came to us and presented a concept that we were comfortable with. And Alex mentioned data, but it's not the data. It's not the data that customers sent to us. It's the data of the performance of the system based on logs. You already had that and the metrics. ThetaRay had a I want to say, it's not a design flaw, but there's no traces that are built into the system for many reasons, some of them architectural, some of them historical. And Logz were able to produce those based on reasoning and detection within the logs themselves. So that already took us one step forward. And then the idea and the concept that Miles and David presented with the comfort that someone is paying attention to what you're telling them. We explained it, I think, from day one that there is a barrier where we can let Logz step into our ecosystem. We were comfortable with where we were. We didn't want to challenge it. Where we were were that there are Logz agents or Logz pods that are basically collecting our logs and metrics and sending them to Logz and stored there. And then anything that we need to do, we get the alerts from there. We get the logs from there. We do the troubleshooting with it. Everyone was very comfortable with it already. And that includes both R and D and the operation side for support and beta teams within Thetaray. Crossing that boundary would have been a lot more challenging because you're letting a vendor in something which is the holiest of ally for us. Okay? And we we take a lot of measures controlling DLP. We take a lot of measures controlling PII data. It's something which is built into our culture. Now that's a challenge because I know that David and Miles would have been happy to be right there with the with the data, with the systems Yeah. With the base code, with everything, and probably do a much faster job. That's the challenge that we gave. And Log said, okay. Yeah. No. No. We can do it. No worries. Now we'll do this. We'll do that. They came up with ideas about how to better the prompts. They came up with the idea. Even the basic desktop application that was written in no time to allow us to start feeding into the model and into the agent and to start to do our work before we were ready with our own part on the other side. Even the patience and the wholesomeness to allow us to do our own side based on examples that we saw that Logz were doing. Okay? So I wanna say that we were very innovative and very creative in on our own. We were working on the design six months before that or eight months before we went live, but there are ideas that we, I wanna say, quote unquote, plagiarized. But not not without asking first. It's called idiocy. I know. Not without asking first and getting permission. And and and we explain basically that what it is that we were doing. We're basically taking, okay. This is the best practice here. Okay. We're okay. Logz's okay with us doing it. So, okay, let's do it because this is what works, and this is what makes the solution works for both of them. We still feed into the model, and then I think that we are now working on closing that feedback loop. Yeah. On feedback. I will discuss this later. Yeah. It's working. And that is what the core basically decision kept on going. We assessed the question of whether to continue with blogs, I think, in every phase that we took to understand, are we still getting the benefit or are we working to for that? Is it working for us? The answer was always on those phases that we are continuing, and the success shows. I think part of the success is that we work on this together. Yes. And I think this is the main point. Okay? Because we didn't feel like, you know, customer, vendor, whatever, it was a real partnership. Because for us, it's new. For you, it was new. I think because we built it together, I think this is why it was such a great success. Okay? We still have a way to go. Okay? By the way, my appetite is much bigger, as you know, and I want to be much more aggressive there. And we're probably going to discuss it in a few minutes. But for us, the fact that we work on that together, where you understand us, we understand you, and this is how we succeeded together on that. Cool. I must ask, why did you try or did you try building this yourself before launch? So we we tried. Okay? As Yossi mentioned, we are very tied on the resources. Okay? It's very hard for us to invest now and to build it from scratch. Okay? It's always great that they have a partner that already have ideas, solutions, and can we can embrace very quickly and to move to production. Basically, when you have it it always a decision buy versus, you know, build. Okay? And on this case, I think for us, it was very natural to go with the buy. Okay? Because, you know, it's already there. Logz are there. We have the data. You have the data. You can understand exactly what's going on in the system. And by the way, we were also very impressed by the solution that you presented a few months ago before, you know, we got to this partnership. Thank you. Yeah. By the way, there is something like, again, given I have the angle from other customers, and I think that AI is cool. Building stuff yourself is cool. And we do see customers, and it's not fair, coming to us and say, hey. Why don't I build it myself? And, you know, the the the everything is progressing very fast. So in six months, it it seems like endless. But we see customers doing the 80% themselves because it's cool and it's fun, and then they're stuck with the other 20%. And this takes time, and then someone needs to do the maintenance, and there's bugs, and there's the things. And so first of all, I'm happy to, obviously, to be a partner of, like, Logz. It's happy to be a partner of ThetaRay. But I think that from our, of course, subjective angle, you did a did a good move because now, like, you early on understood that this is not a click, click, click, everything works, and that's it. It's a process. It's a strategic process. And I think that for the two sides, this was a very I think, like in every area in life, you need an expert. Okay? Yeah. We have the same issue that you just described also with our customers. And as I mentioned, we have our own NGINX solution called RAY. And we are expert in non national crime detection, okay? And we want to push our customers to use RAY, to embrace it, etcetera. Also, our customers tried to build it by themselves and they couldn't because we have much more experience, we have much more knowledge and similar to you, okay? So from our perspective, we wanted to go with somebody who can give us additional aspects that we never thought before. Okay. You have much more visibility to other customers. You own customers. You understand the business. You understand observability. You understand how it works. And from at least from my perspective, I always prefer to go with an expert than try to build by yourself. By the way, we are building stuff by ourselves. Yeah. Yeah. Yeah. Yeah. We we have a lot of AI initiative across companies. We have RAY and also internal tools that are part of NovoProject. But when we're touching production, when we are, as Yossi mentioned, the holy of holiness, You want to be with the most solid solution that is in the market, okay? I don't want to make experience on production. I can make experience on other parts of our, let's say, workflows, etcetera, but not on production. In production, we need a leader. We need somebody with knowledge, and we need an expert. And I think you have, like, three checks on what Jill Thank just you. So maybe that's a tip to our audience. Like, if you're considering building it yourself, focus on your category, on your domain. Can you do it? 80% for sure. But I want to say where the boundary is Yeah. In that matter. Processing the data and doing the reasoning is something that you can go and test a lot of models and a lot of things. And if you have the time and the resources, you will train it. And eventually, we'll get something that you can work with and all that. However, if you're going to put something that is automatic, that is going to disrupt the structure of your services, focus on the services side and let the reasoning be done by the part that manages your data. Whoever owns the data owns the goal for that matter. That's usually the case for us. And it's also, as we said, that's the rule, as Alex said about our own products. We focused on implementing our ways of working, of taking auto NOC and making it replace the personal NOC, the student that was there for twenty fourseven. We implemented in the tools and the ecosystem and the different variety of system that we use, whether it would be for ticketing or alerting or on call management and everything in between. We focused on that development, on that infrastructure, and then accommodating the place for the part that Orion came and basically gave us the reasoning and the answer and the path to resolve or the path to report and to do whatever the playbooks are And telling this is where my recommendation would be because I went through this via De La Rosa to get that. Yeah. Okay? So for the audience, ORIONIQ mentioned also ORIONIQ, this is the name of our agentic observability platform of Logz. So maybe we should move forward. How did you decide what the first Alex, how did you decide the first agent would be? So I think we experimented. We try, you know, we first of all understand where our main issues are. Okay? We tried to gather the data, understand, okay, should be our focus. And based on data, we decided well, what like, in the regular process, you prioritize based on data. You prioritize based on the KPIs that you build. And for us, it was very easy to understand. Okay, first of all, we need the first agent is detection. First of we need to understand what's going on in the system. Okay? Then we can understand afterwards, Okay, we detect the issue. The second phase is to do the RCA. Okay? The third phase, and we're still not there yet, but we're going to be there, as you know, is try to resolve it not only by humans, but also some of the processes by the machine. Okay? And I think this is the stage where we are. We still want to see how the platform is performing. We still want to understand that we are in a good shape. But I think the next step, and for me it's super important, is also try to resolve some of the issues automatically. Automatically, yeah. Okay? Because there's issues that are more of the same. It's somebody that needs to go and do select, select, select and whatever and press submit. Okay? For that, we don't need a human. I think human needs to focus on the smart stuff, okay, on the stuff that can give us additional value. Okay? And the agents and the machines need to do the dirty work. The mundane. The something that is manual, that is not needed anymore. And and as you mentioned, okay, this revolution is ongoing. And you don't know what happened within one month or two months. I don't think any of us a year ago when we started was able to imagine what's gonna happen right now. We still cannot imagine what's gonna happen in six or seven or eight months from now. I think every month they're passing by, better models coming, better solutions in the market. And I think we want to be make sure that we are up to date. Okay? And we have the best proactive support system in the world. Okay? And we're gonna be there. Okay? I'm pretty sure that with your support, okay, we're we're gonna achieve this very fast. Cool. I wanna take this a little bit too historical because when we started the scoping for what will be the first thing that we're going to implement, we decided to tackle the p one critical issues. That was very naive and very hubris. And the fact that you came to us in the middle of the process make us think a second time and a third time. Then I think that we went with the DAGs, right? We went with a mundane processing issue that we had in the point of interface between ourselves and the customer, something that happens every day and generates a lot of data and generates a lot of noise. And that was the smart choice. And that smart choice was taken because we were led to that water hole to drink the water. So thank you for that. Other than that, yes, there's scoping coming for another phase, we're probably going to do amazing things. I asked Claude if we are the first, and we are within the top five and first in the country and in the market. From our customers, you're definitely the most advanced. Yeah. Thank you. Maybe, guys, you want to share, once you started to use the Ryan IQ agents, how did the NOC day to day work has changed? I think from our perspective, as I mentioned, and I will repeat it and repeat it again and again, we move from reactive to proactive mode, okay? And for us, this is super important. Our people right now are focusing on the resolution and not on the detection. And this is the major shift in the mindset of all the support organization. We have the solution for the detection agents. They're doing a great job. Now we want to make sure that we are resolving issues much faster. We provide a better support to our customers. And this is where our focus are on the daily basis, where we measure it, where we're understanding that, and this is where we want to be. Okay? We don't want to be in the boring world of detection. We want to be in the world of resolution as much as fast as possible. Okay? And we're going to get there together with you. Of course. And when the team transitioned into the new roles, you mentioned, you spoke about it, Yossi, earlier that people moved up to different positions. How did the AI became as part of the workflow of that? So first of all, I will take it for a second. We have Yossi, okay? Yossi basically is the one who actually established the support in Sound ThetaRay and they also established the NOC. Okay. So he has a lot of knowledge. When we talk and what's he what is next? It was obvious to me, okay, that we need this guy with all this brand and knowledge that has inside his head to now be the AI leader or AI champion inside our support organization. Okay? And I think this is like this is the new role that we created. Okay? Yossi now has all the visibility to all the workflows, to everything that is done by the support. Okay? He understands exactly what's going on, and he is the one who guiding the team, okay, to build AI solution. Beside your solution, of course, that we're working on NOC, all our SRE team now creating internal tools using AI in order to improve our, let's say, services. And Yossi is the key person there. We guide the team. He understands exactly. He has all the three sixty view. He understands exactly what's going on. Okay? So for your question, this is the first role that we're creating during this process. I think on the other things, correct me if I'm wrong, the tier one now a real tier one. Okay? They are not looking at the screens. They're trying to understand what happened, okay, what needs to be done and how to resolve it very fast and quickly with the other team in the company, okay? And I think this is the major shift also that happened to us. Understood. There's something which I wanted to raise is that and I think again, like, thinking about what our audience is interested is, how does an implementation project look like? We Wow. It's hours and hours and hours of thinking and drawing on the board and deleting and thinking and rewriting it again and trying to find the right people from your team to participate in those think tanks. I think we had at least within a span of a year or or even the six months, somewhere between six months and a year, before we had like three or four off-site sessions that we did within ourselves and with partners within the company that participate within the support process, so escalation points, teams, and all that, to figure out what everyone wants from it, how should it behave, what it is that we're missing and we're not looking at it, taking points of views from different places. Our team is spread up on three continents. Okay? We're in Israel. We're in Spain. We're in Latin America. Those are different, not just geographical location, those are also point of views of people's day to day life. And that is important. So you listen to everyone, you gather everything, and that is just the scoping. When you get to the implementation itself, Eventually, beside the fact that I think the fact that we use existing Nock engineers and Pelek who are a team leader of Nock, okay, was heavily involved in this project, he understand exactly how Nock is working, and it was very easy to gather the requirements. Okay? So I think it's before we move to the implementation, I think the first thing now, especially with AI, is to create a solid list of requirements. I think this is the key point. Once you have that, then as I mentioned, execution now is easy. Okay? It's not rocket science anymore. But to gather all the requirements to understand how the system needs to be will behave, It's this is the most critical part of the implementation process. So from our perspective, the scoping and the design, we invested, as you know, a lot on that. Okay. It took us months or two, I think. No. Scoping was a lot. Even more, I think. Six months. But we want to make sure that we are putting something solid in the production. Okay? And that will gonna help us and not gonna create us more issues. Okay? And I think for us, this was the most important thing of the implementation process. Yes. So take the mix of the people that were there. We were the head of support. We were myself, Perig, who leads the tier one, Eamon, who used to be a tier one but progressed to tier two about a year before that. So he has point of views from both locations. And this is Yossi, this is, like, before we met? Or This is in the months before your offer came, and we started to do the scoping. And then somewhere in January, your offer was injected into the process. I think we met in January. We started to draw that circle in the middle as And an why? Because, again, we circled back to all the process, all the things that we wanted to do. Eventually, I don't want say it just to brag, but the development. And we used wide coding. We used automatic tools. Okay? We tried several platforms, by the way, based on the licenses that we had and whatnot that we have. And then all of that investigation, all of that research that we did took us into a focused development process and training process that was about three months, okay, to go into a soft launch at the end of June, middle July, and then to go to a full launch at the beginning of August and then eventually declaring it fully on, full production, which we're doing now in this September. That concentrated effort came to be because we knew exactly what it is that we want to do. And we had to, you know, correct posts in the middle like always, but we had a solid idea of what it is that we want. We had a solid idea of what would be the method. We started from the core, and that is the important thing to, you know, to tell the listeners and people on this podcast. We started from the core, and that is to train the reasoning model with our playbooks and with the process and to start to feed back into it. We did it with originally with the Orion desktop tools. And then later on, we did it with our own development, which is the Orchestrator. That's the Autonox heart. Everything else was another layer and another layer and another layer, and we'll keep adding them. We already have a collection of skills that are unique to us and to things that we need. There's one that I owe Alex, like, for a couple of weeks, but today is probably going to see some of it. We do mistakes, by the way. Okay. We we did mistakes. We learned through actual exercising. That's okay. We don't do mistakes with the detection. We do mistakes with how do we manage our steps and recommendation and all of that. We get automatic phone calls when something is wrong. Okay. I get wake up at least twice a week because some logic there decided to wake me up. Four false, not positive, and then we go back and fix that. So it's constantly happening. And the fact that we started from the core of the reasoning and then moved out into how do we manage the process, how do we take it into our day to day, how does it communicate with the ticketing, with the alerting, with all of that, That was, I think, one of the key success points that we implemented into it. So that's what the implementation looks like. And it's a lot of hours. It's a few nervous breakdowns, a few dramas. You know? Natural. Every is legal life. I think that one of the things I like, from the side looking at Thetaray, you came realistically to the process. Like, you knew it's going to take time. You knew you would need to invest even more than, like, you you typically do. Like I don't want to judge, But ThetaRay is a company culturally, okay, was never a company of big buzzwords. Okay? With the formal leaders, with anyone that came and go and had a Sonoco, we're not a buzzwords and we're not a sorry for the French BS company. Okay? We don't do the big events. Okay? We don't do, like, a full party of on the other songs. It's not important for us, sorry for that, to be in the main tower in the center of Tel Aviv. Okay. We are a company that started at Har Hotspur in Jerusalem and moved to the back end of Hodesharon. Okay? We always joke about the fact that it's hard to recruit people East Of Route 4. Okay? Yeah. That's the company. It's a tough less company. Yeah. Okay? This is why we have a very small margin of people. This is why we do a lot with very little resources. Okay? It's part of the, I wanna say a little bit spontane essence of the company, and we're proud of it because it makes people think outside of the box constantly. It brings ideas. And the fact that we have AI today to bring those ideas to life, that is a perfect combination of the two. Okay. This is why, by the way, a lot of the success that you see for now for the Otonok, I can take some of the credit for it, but I have to give a lot of the credit of the team who worked with me because every idea that they had, they could they were able to produce, put it in place, test it, and have the integrity to scrap it where it didn't work and to start over. Yeah. And they did. So so, like, I think that the culture that you mentioned, look look, I know by now, I feel I know a little bit of your culture. I think that, like, for companies that want to undergo this task of automating, they need to understand that this is, first of all Move to good culture. I would Move to Hodeshawn. Move to Hodeshawn. It it requires a certain culture, but it's an effort. Yes. Strategic changes with opening doing your email is easier with the chat Doesn't require any any any change in the company. But to go what you did requires, like, understanding that it takes time. It requires, like, leadership support, not for a week, for months. And I think that the benefits are clear, but it takes months. It's not if someone thinks that they will be able to automate their NOC in a week Nope. They're wrong. No. They're wrong. This won't happen. I think the most critical part is not too afraid. Okay? Yes. And I think once you put your fears outside and you try, okay, you find amazing stuff. And we discover a lot of amazing stuff during this project. We also improve our internal processes. And I think this is also a great win for us. Okay? So I think don't be afraid. Okay? If start small. Okay? And jump bigger and bigger and you get the appetite. You will see the progress. It's like today, you see the progress day by day. Okay? Basically, we see how we are becoming much better, how we are becoming much more proactive. And I think start the project, gather, as I mentioned, do the scoping first, do the design first. These are the most critical parts. Okay? Once you have everything, okay, you can start towards the implementation, but don't be afraid. Okay? For us, production is the holy of the holiness. But again, we are doing it very gradually. We test everything. We tune stuff, as Yossi mentioned. Okay? We understand we have let's say we have discovered false positive. Okay. We understand what not to do now. Okay. And we're fixing it on the fly. So for us, it's a great success story. I want to add to it. The moment that we launched the work and those three months that we started the development and we shared it with the rest of the team globally, it raised the appetite. Okay? There was a little a sense of greed. People want to be in. Okay? They were willing to step on another. I want in. And we didn't because it would have made a mess. Okay? We kept it close and tight. But it doesn't mean that the other people lost their appetite within the team. Yeah. When we finished and we started to launch, you all of a sudden saw that waiting for you at the back end, everybody prepared their own So tool for the pain point that they the barrier was broken. You can do stuff for production that will actually do stuff for production, you can build your own tools for it. That barrier, that that fear was broken now, and we benefited it not just with having the autonomic and the Orion reasoning within it, but also an entire ecosystem that people started to build to improve their day to day and to improve customer service pain points that they had. And we see it. We close cases faster. We respond to them with more details. We communicate somewhat better. We have to improve. And now we know that. And we are able to think now about a scope for another phase, okay, with, I think, that the bucket is the basket there is about three times as big as phase one at the moment, and we need to, you know, come come realistic to that. You know how I I know the people are using? They asked me to increase the subscription to Claude and to ring. Yeah. And then so I'm paying a lot of money right now. And topic just called to say thank you. Yeah. So, like, we already in the I think some of our engineers already on the $100 per month subscription. And now even some of them on 200. So I will not tell you, not you. But some of them basically started to become addicted to it. We see people that working during the nights, during the weekends in order to build those tools because and by the way, I'm doing myself. Okay? I build, for example, some of the pages in our website. Okay? Because it was a pressure to do something fast. And I together with our marketing, amazing marketing team, we built it together. So it's super easy now and the people really want to embrace it and to use it because I think the smart people understand that if you're not there, you're not gonna be part of this new market that is gonna be established. So everybody wants a piece of it and everybody wants to be part of it. And I think it's great because basically it will make us much better. I think you've touched about the next phase. I want to ask you both, like how do you see the AI NOC as So point as I the mentioned, we're now the goal and I will repeat it again and again is to become proactive, okay? We're now detecting. We understand. We have the analysis what happened. Now we need to think together how to fix stuff automatically. That's what I'm fixing. Okay? Yes. Still should be a human in the loop that will review what the agent want to fix, okay, and to do it, you know. But the resolution time needs to be much less than we have right now. Okay? We need to understand that most of the, I think, issues that we have, it's not, like, it's not something very hard to fix. Some of them, it just, you know, easier to increase some port or, I don't know, to increase some memory or whatever. And for us, we don't want to be in the position that for every issue, we need to put a person on the ground in order to fix it immediately. And we need to move now to a smarter resolution that we don't harm anything on the production level. But still, we are fixing some of the issues automatically as much as we can. Understood. I'll give you that it's you're not letting a reasoning tool reason itself how it's going to fix it. You provided with ready made solutions, Okay, and allowing them to choose how to try to resolve it. And by that, I mean is that you separate the resolution from the reasoning and you're giving it an action, which is a resolution. And the action can be a Jenkins job or a ready made script that the engine doesn't knows only how to apply to it. So it applies the variables and and triggers it. Yeah. On the repetitive stuff, absolutely. Yeah. Absolutely. That makes complete sense. Otherwise, if you're letting the thinking think that it can resolve itself, that's the portion that we are putting, at least in the scoping, simulate, but don't apply. Not yet. Yeah. Understood. That's in the core of next phase. But other parts of the next phase is that's the core. That's the biggest investment that we're going to have. Otherwise, around it, time wasters. So putting down an RCA in a language that, let's say, not everyone has, okay, the level of written language that you sometimes need to submit to a bank which is a foreign in a foreign country, English speaking country or Spanish speaking country, that's a task that can easily drink away four hours away from So someone's no, that's a time waster. You put the information in just common language, whether it would be Hebrew or English, and you let it produce the dossier for you with the evidence that it already has. Configurations, maintenance, a lot of those things, things that are time wasters, those are going to go into the second phase because it clears up the time of our SREs and our team to actually engage with the customers, to investigate further, to do some better discovery as part of the troubleshooting because it's very important to us, And to invest internally because we have a lot of internal initiatives within support and generally within ThetaRay that we want our people to be part of it, and we want to clear out the time for them. Understood. So we're getting to the end of the hour. Assume there's a company thinking about automating their NOC. Guys, maybe you can give them, like, two, three tips. I think I mentioned What do she do? Yeah. So I think I mentioned So lose fear. Put the fear out of this. Second thing is to do scoping, like review all your playbooks, understand exactly what's going on, when you react, when you don't react. Okay? Like, this is the most, as I mentioned, this is the most critical part of this entire journey, okay? You need to do the proper scoping, you need to design your system, and you need to understand exactly how your system is working today, as Yossi mentioned. For us, we had Peleg. Okay? He was with us for a few years, and he has the brain and the knowledge exactly what's going on in the NOC. So and as you mentioned also before, invest. I think you need to make sure that your management and leadership are supporting this move. Okay? And they need to understand that they also need to invest for a couple of months, if not more than that. Okay? Because it's like a journey. Okay? Yes. We launched it together. Now we're in the process of tuning the alerts and now that we understand what's going on. The next phase, already discussed about it. Okay? But first of we will not move to the next phase until we're gonna finish all the tuning and make sure that what we already launched is working properly. So for me, fear out scoping in, okay, and the support and the of the leadership and the management team in order for these processes to take place. Understood. Talk about ROI. Oh, wow. So to be honest, I think for as Yossi mentioned, from our let's say, the ROI is and it's not only how many money we save for me. Okay? It's how better we become as a support organization. To to your customers, you mean? Yeah. Absolutely. Because, you know, we didn't lose too many jobs. Okay? We we just shifted those 18 students to three FTEs. So the best arrive from me right now is not only on the resources side, it's more how better we become as a support engineer, okay, and our support organization is much more proactive because eventually we have too many issues and we are not communicating well enough, okay, to our customers. We don't explain exactly what happened. If we don't resolve the issues fast, we're losing money, okay? And for me, this is the most critical part. Okay? We need to make sure that we are providing much better service to our customers. Yes, we save some money. Okay? Yeah. No. There's nothing bad with that. Right? We save some money on the resources. You're right. But it's not a huge amount. Okay? From our perspective, the most critical part is that to provide a much better service to our customers. I would pinpoint on to that the fact that the support has now become a sell point and not a negotiation point. Oh, cool. So once you tell them, yeah. I have this tool that I can react within a couple of minutes, and I can detect, and I can connect it to your monitoring tools, and I can send you x, y, and z. And then there's a robot that's gonna be there twenty four seven and is going to answer, and then you're going you're covered, we're covered, and it will wake me up. It will wake you up if you want. Don't worry. You have something that you can present as a process and as a selling point and not as a negotiation point. Understood. Okay. So we discussed about avoidance, about avoiding the penalties or avoiding the bad service or losing money while you're down, but then you can also show those benefits, And we can already see them in terms of feedback from the customer. And had a meeting with our sales team. It's an annual visit they had. They usually are greatest critics as always. That's Okay. We don't think that they are that much. But it's a rivalry within Tatore always. To this year. Get to walk with them. Yeah. It's a rivalry that we have, but this time I sat down in a room and heard at least two people praise the improvement, the improvements that they saw in the last couple of months. They said that we answer faster, we close faster, we're more clear to that, we still have a way to go. But they saw the improvements, and they felt it, and the customers gave the evidence to that. And we didn't tell the customers too much about the auto NOC, okay? They're not really aware of it. They know that there is a 20 fourseven coverage for anything that happens that is a critical issue. They know that they are getting answers. And they're saying that service has become better within a very short while. So happy to. Wow. Yeah. So I think we just ran out of time. So Alex, Yossi, Gali, it's been a pleasure. Can't wait to see what we build together in the future. Thank you. Thank you. Thank you for having us. Thank you so much. Bye bye. Thank you. Bye bye. Bye bye.
At a certain scale, alert volume stops being a people problem and becomes an architecture problem. ThetaRay reached that point and rebuilt their NOC operations around AI-driven triage using Open360 AI and OrionIQ agents to address it.
In this fireside chat, Alex Glushenkov and Yossi Cohen tell that story in their own words.
What you’ll come away with:
Speakers:
ThetaRay:
Alex Glushenkov, SVP Customer Engineering
Yossi Cohen, Senior AI and Operations Engineer
Logz.io:
David Lotan Bolotnikoff, VP Product
Gali Mounis, Enterprise Account Manager