AWS Summit 2024

Posted on July 10, 2024 by Jeanne Boyarsky

I went to AWS Summit New York today, a free one day conference. It’s the first time I’ve gone. I didn’t live blog but am writing a summary post of my day after the fact.

Overview

AWS spent a ton of money on this event. They rented out all or most of the Javits Center in NYC (this is where NYC Comic Con is held). They gave coffee/soft drinks and even free lunch. They also spent a lot of money on security. For cause. There were protesters right outside the front door.

I tried to experience the major parts of the event.

Expo

The exhibit hall was large on the third floor with lots of vendors related to cloud. There were also some fun activities like a drone and toy car racing. Lots of space for sitting/networking.

There were also some stages in the expo for shorter (15-30 minute talks). They had headphones for people who couldn’t filter out the background noise of the expo. It was nice because you could flit by and see if you were interested. I listened to some pieces of cert/education talks and a full one from Elastic on LLMs and summarizing security incidents

Breakouts

There were lots of one hour breakout sessions on the first floor. I went to two customer success stories (Venmo and Fannie Mae). It was dark in the breakout rooms. Like most places have for keynotes

Learning highlights for Venmo

Key strategies

Distribute load to maximize processing throughput
Use event based systems for anything not in critical path

Other notes

Django app. Used Celery for async work, reader db instanes for queries that can use
Then added DynamoDB, MongoDB, OpenSeach Service, data lake, microservics, Cassandra (for microservices), Kafka
Split Mysql into Auroa MySQL comatible secondary ySQL and Analytics MySQ databases

Social feed data migration

Transactions visible, high traffic because home screen
Every transaction geerations a feed story along wih certain profile operations
3.6TB of data, 5.6 bllion entries
Since digit lateny on data retrieval
90% of memory usage
switched to DynamoDB due to cost (90% less), performance (equivlanet), managed servie, data encrption at rest, integration with other AWS offerings
Migrated via backfill followed by dual writes. Let verify performne under pro load and confirm data consistent. Then started ramping reads on new database. Started with 1% reading from new DynaeoDB. Finally cut off writes to ol MongoDB

Offloadng transaction history

For each payment put message on Kafka queue and write to Cassandra via microservice. Implemented as best effort write Needed to guarantee 100% of data so could move over use caes taht required full fidelity data
Switch to write ahead log – write log essage saying intend to peror action and store in DynaeoDB Then proess transaction/pblish essage. FInally, delete inteded action message ow that completed. Background process looks for pendin messages
Asyc payment processin using Kinesis
Problem batches huge and inconsistent for credit car sage, delays, outage costly, can’t send 500 error/need to reconcile, not a way to replay transactions internally
Added Kineis Data Stream via think wrapper to put mesage on strea and ackowledge success to upstrea. From KInsis, have consumers/lambda procss. Also usig Auora, DocumentDB, ElastiCache, DynamoDB and SQS

Key learnings for Fannie Mae

data science research

compared research vs deveopment – ex: research has poc, live prod data, latest tools/patterns
pilars of platform:
data access – prod data, data usage contracts
governance – control by business, not tech, autoamted integration with governace
operationalization – testing, validation, Ci/CD
data science controls
register research activities in CMDB so can provision/tag resources. Automated provisioning, strealined architect review process
Data access.sharing contracts, perissions, ingress/egress rules, sensitive data protection rules
Cde deployment and change managment CI/CD, scanning
Data science platform architecture
code/image repo
pblic data endpoints
code/package library
read only access to enterprise data lake
research envs –
collaboration – just in time access – read only access to prod enterprise data lake. results an’t be shared; considered dev
validation – testing/shakeot – still read only
operaiton – headless execution/- now can write to prod, create reports and share exterally
data access JIT (just in time). Fannie Mae has a patent on this
request access to data. could be from many data sources
JIT access engine checks against coarse grained contracts
Then goes to policy manager to check fine graine access controls. Use UI to create rules. creates new role dynamically so can use token to access

Building a generative AI use case

Used Anthropic’s Claude 3 Sonnet via Amazon Bedrock and Aazon Neptune (graph db)
A lot of analysis of unstructured documents, average of 5 hours per doc and 8K dos per year
Deep Insight for LLM driven knowledge extraction. Uses ontology (schema( an LLM t generate knowledge graphs. Human in the loop to validate Then knowledge utiilization step to use natural language via a chatbot
taxonomy – linear top down hierarchy. Ontoogy – interconnected network representation
Disambigution important to avoid duplication
graph database
reduces risk of hallucinations because more context
two types –
Property Graph (Apache Tinkerpop) . Query with Gremlin or Cypher
RDF Graph (from W3C). query with SPARQL
extraction uses Bedrock, fargate, lambda, neptune, s3
utilization uses – bedrock, fargate, neptune and a chatbot
also uses LangChain – Neptune Open Cypher QA chain (converts natural langague queries into Cyper so can do query( and Amazon OpenSearch
challenges
pick onthology framework – Chose Turtle (Terse RF Triple Language for reeasability/ease of reading
find best way to chunk. Chose at sections so handle complex tables btter
Picking graph type. Chose property graph due to better OSS framework support
Amazon Kendra (enterprise search( did not integrate with Amazon Neptune. Used LangChain’s NeptuneOpenCypher QA Chain instea

Chalk Talks

Chalk talks were also on the first floor. They were also an hour but had less prepared content. The one I went to had 20 minutes of talking/demos. Most of the time was Q&A or discussion. They had a whiteboard with a camera to show what was on it so the speakers could write/draw real time. This meant one projected screen was the computer and one was the physical whiteboard.

Learning highlights

gen customers what to know what model to use, how to move quickly and how keep data secure/private
Bedrock provides foundational models via single API, customize model, RAG (Retrieval Augmented Generation), agents for multi step tasks, security/privacy/safety
Models include – amazon’s models, anthorpic,, cohere, meta, etc. ANd lots of variants/versions of each.
Two use cases: observability of generative AI itself, using gen AI to help with observability
gather metrics – ex: number tokens used for input/output
collected metadata/requests/responses so understand how customers use
governance/controls/guardrails
Cloudwatch – analyze inovcation logs, protect sensitve date, real time metrics and alarms (Ex: more latency on different version of claude), single pane of glass/dashboard
recorded demo #1 (while video was recorded, he narrated live. also paused periodically to say more
can send model invocation logs to either s3 (if using other loggiing system) or cloudwatch

Builder Sessions

Also on the first floor, these were small group labs. I went to one on Amazon Q. They had 4 areas on the room with 10 chairs each. An instructor from AWS was allocated to each group. After a short intro, the instructor helped anyone stuck and answered questions. This was great.

The lab had an access code good for three hours so you continue a little longer if you wanted. In theory, there was separate wifi for the lab but it didn’t work. The main conference wifi was fine though.

Learning highlights

Amazon Q Developer has a free and paid version.
The paid version promises not to learn from your data, It’s licensed per person but only billed if the developer uses in a month.
IDE integration for VS Code and IntelliJ.
Chat bar. Often gives sources/links. From 2023 for public internet. RAG for Amazon so more recent
Can explain code, refactor code, fix code and migrate to later version of Java. Can also write a plan for writing code and write code (with some errors)
Code Whisperer was folded into Q
It was slow, but I was on a conference network

Main dev activities

planning – docs, examples, deisgn
- creating = generate cpde,amage omfra
- test amd secure – test cases, scan for security vulnerabiliteies
- operate – identify and mitigate code issues, monitor performance and efficiencey
- maintenance and modernization – modernize and update old code languages and dependencies

Amazon Q Developer tries to help with all phases

plan – explain code with conversational coding (chatbot)
create – inline code complete, conversational coding
test/secure – unit test generation, OWASP top 10 security scanning
operate – debug/optimize code with conversational coding
maintenance and modernatization update code with agent from legacy

Keynote

The keynote was in a big room that wouldn’t fit everyone. They also used all the breakout rooms as overflow and streamed to the stages in the expo. I like that as it was easy to eat and listen. Or talk to the vendors and listen to parts. Or not.

[devnexus 2024] refactoring after fowler: some large refactoring patterns

Posted on April 11, 2024 by Jeanne Boyarsky

Speakers: Aaron McClennen & M. Jeff Wilson

For more, see the 2024 DevNexus Blog Table of Contents

This talk inspired by books:

Fowler’s Refactoring
Kerievsky Refactoring to Patterns
Gamma (Gang of Four)
Feather’s working effectively with legacy code

Refactoring

Restructuring code without changing behavior
Purpose of computer language is to tell other programmers what to do. The computer uses ones and zeros
Make it work, make it right, make it fast
If put it off, will never have time
Code smallers to remove – showed screen of examples
Refactor to reduce WTF/s minute rate in a code review
SAFe 11. 4 “refactor to support the new behavior of the code” – one of the built in quality practices
Do when need to change code, bug hard to fix, need to reduce tech debt, etc

Staying safe

Want high test coverage
Start small
Proceed incrementally
Test after each change. Undo last change if fails
Use tools like Veracode and Sonar to find code that needs changing

Simple Example

Need to add a flag.
Introduce Parameter from Fowler.
Showed how adding another flag is trivial

Planning a refactoring

Think about like planning a trip.
Decide trip is necessary – overcome inertia
Scratch refactoring from Feathers – do a refactoring to get familiar with code and revert when done. Helps figure out what to do, Think of as “exploratory refactoring”
Select a destination – Understand what would like it to look like
High level refactoring plan
Refining the route – more details
Make it so

Example

Showed method with two parameters – a dto and interfaces
Showed Template Method Pattern
Plan make a base class, turn implementations into subclasses, remove interfaces and stop parameter passing
Refine plan: map to concrete steps like drop the interface and stop using the interface. 13 steps to do the four higher level steps

My take

The intro felt very long. Would have been nice to see if audience needed an intro to refactoring. First example at 20 minute mark (for the Fowler example) and 23 minute mark for first mention of planning for large refactorings. I was speaking after and left early to get ready. So I suspect I missed some of the best parts. I was expecting more of it to be about patterns. The part I saw was too easy for me.

[devnexus 2024] ai proof your career with software architecture

Posted on April 11, 2024 by Jeanne Boyarsky

Speaker: Kelly Morrison

For more, see the 2024 DevNexus Blog Table of Contents

HIstory

Fairly recent. GPT created in 2018. Number parameters increasing exponentially
Microsoft CoPilot released in 2021. Uses Codex; a specialized model off GPT3 for creating code. Trained on billions of lines of GitHub code and can learn from a local code base
Amazon released CodeWhisperer in 2022. Can generate code for 15 languages. Specialized for AWS Code Deployment

Basic Example

Asked ChatGPT to write a Java 17 Spring boot rest API for stats in a MongoDB with JUnit 5 tests cases for the most common cases
Looks impressive on first pass, but then find problems
Hard coded info
Used Lombok instead of Java 17 records
Code doesn’t compile

Complicated Example

Asked ChatGPT to write an entire enterprise app for selling over 10K crafts with a whole bunch of requirements like OpenID, Sarbanes Oxley, etc
Didn’t try. Instead came back with a list of things to consider in terms of requirements

What AI can/can’t do

Can do “Ground level” work.
Still need humans for large orchestratoin – ex: architects
Can do more self without junior devs
Garbage in, garbage out. Trained on public code in GitHub. Not all good/correct. Some obsolete.
Humans better at changing frameworks, working with CSS (does it look nice), major architectural changes, understanding impact of code when requirements change

Hallucinations

Doesn’t understand. Asks as mime/mimic/parrot
If can’t find answer, will give answer that looks like what you want even if made up. Example where made up up a kubectl option
Not enough training data on new languages/technologies. More hallucinations when less training data
Mojo created May 2023. Likely to get Python examples if ask for Mojo. However, it is a subset of Python with some extra things

Security Concerns

Learns from what you enter so can leak data
Almost impossible to remove something in a LLM. ex: passwords, intellectual propery, trade secrets
Some companies forbid using these models or require anonymous air gapped use. Translate something innocuous into what actually want

Debugging

Can human understand AI generated code well enough to debug
GPT and Copilot can sometimes debug code, but have to worry about security

Pushback

Law – ChatGPT made up cases
Hollywood strike – copying old plots/scripts/characters
Unclear if generated output can be copyrighted. For now, not copyrightable but could change.
Some software is too important to risk hallucinations 0 ex: plane, car (although Telsa getting there), pacemakers, spacecraft, satellites
Lack of context – other software at compnay, standards, reuse, why use certain technologies, securities
Lack of creativity – need to determine problem to solve or new approaches

What AI does well

Low level code gen (REST APIs, config, database access)
Code optimization
Greenfield development
Generateing docs or tests
Basically the kin of tasks you hand off to a junior developer [I disagree that some of these are things you hand off]

Career Advice

Focus on architecture, not code
Don’t just learn a langauge or framework.
Learn which langauges are best in different situations
Learn common idioms
Look at pricing, availability of libraries and programmers
Learn which architectures should be implemented in different languages
Learn how to create great prompts for code generation
Learn how to understand, follow, test, and debug AI generated code

Book recommendations

Building Evolutionary Architectures
Domain Driven Design
Fundamentals of Software Archicture
Head First Software Architecture

More skills

Types or architecutures – Layered, event driven, microkernel, microservices, space based, client/server, broker, peer to peer, etc
Determine requirements – domain experts don’t know enough about software to specify. Can be bridge between AI and domain experts

Mentoring junior developers

Teach how write high quality prompts.
Remind to ask for security, test cases, docs, design patterns, OWASP checks
Show to spot and deal with hallucinations
Help to understand and debut AI written code
Help learn architecture by explaining why choices made
Ensure code reviews are held
Precommit git hooks to test code
Use AI to help generate unit tests

ArchUnit

archunit.org tests architecuture.
Can add own architecture rules.
ex: never use Java Util Logging or Joda Time
ex: fields should be private/static/final
ex: no field injection
ex: what layers are allowed to call
Can include “Because” reason for each rule
Ensures AI doesn’t sneak in something that goes against conventions

My take

Good examples. I was worried about the omission of “where to senior devs” come from but there were examples like changing frameworks so not entirely ignored. Good examples from the ecosystem as well. Good list of skills to focus on.

Down Home Country Coding With Scott Selikoff and Jeanne Boyarsky

Java/J2EE Software Development and Technology Discussion Blog

Category Archives: Conferences

AWS Summit 2024

Overview

Breakouts

Chalk Talks

Keynote

[devnexus 2024] refactoring after fowler: some large refactoring patterns

[devnexus 2024] ai proof your career with software architecture

Overview

Breakouts

Chalk Talks

Keynote

Share this:

Share this:

Share this: